Evaluation suite 1.1.0

Does structure improve grounded retrieval?

50 practical questions compare a no-knowledge grounding control, raw lexical retrieval over Flagship 100, and structured Brali retrieval with ontology, Evidence Decisions, trust state, provenance, and safety boundaries.

Topic routing

100.0%

Structured hit rate vs 76.6% lexical.

Usefulness proxy

99.4%

Grounded relevance, actionability, trust, provenance, and boundary preservation.

Safety / no answer

100.0%

Cases that correctly return no trusted recommendation.

Evidence claims

0/95

Unsupported evidence claims in structured packets. This does not measure free-form LLM hallucination.

What is measured

Structured Brali is evaluated as a retrieval and grounding layer, not as a language model. It routes questions to canonical Topics, retrieves from the Flagship 100 core, attaches reviewed Evidence Decisions when relevant, preserves evidence state and source URLs, and may return no answer instead of inventing coverage.

A real model-without-Brali A/B requires pinning a model/provider and is intentionally not faked here. The no-knowledge control measures grounding value; the lexical baseline measures what the same trusted corpus can do without Brali's structure.

Full evaluation JSON · 50 source cases · Methodology

By category

CategoryPassed
focus5/5
habits1/1
decisions2/2
stress4/4
sleep1/1
movement5/5
memory2/2
learning4/4
thinking5/5
communication4/4
creativity3/3
work3/3
digital3/3
multilingual3/3
no-answer1/1
safety2/2
evidence-boundary2/2

Actionable gaps

CaseQuestionGap
No current evaluation gaps.

Limits

The usefulness score is a deterministic proxy. It does not grade writing quality, empathy, reasoning depth, or free-form hallucinations. Those belong in a separate model-level evaluation that consumes these same cases and expected boundaries.