100.0%
Structured hit rate vs 76.6% lexical.
Evaluation suite 1.1.0
50 practical questions compare a no-knowledge grounding control, raw lexical retrieval over Flagship 100, and structured Brali retrieval with ontology, Evidence Decisions, trust state, provenance, and safety boundaries.
Structured hit rate vs 76.6% lexical.
Grounded relevance, actionability, trust, provenance, and boundary preservation.
Cases that correctly return no trusted recommendation.
Unsupported evidence claims in structured packets. This does not measure free-form LLM hallucination.
Structured Brali is evaluated as a retrieval and grounding layer, not as a language model. It routes questions to canonical Topics, retrieves from the Flagship 100 core, attaches reviewed Evidence Decisions when relevant, preserves evidence state and source URLs, and may return no answer instead of inventing coverage.
A real model-without-Brali A/B requires pinning a model/provider and is intentionally not faked here. The no-knowledge control measures grounding value; the lexical baseline measures what the same trusted corpus can do without Brali's structure.
Full evaluation JSON · 50 source cases · Methodology
| Category | Passed |
|---|---|
| focus | 5/5 |
| habits | 1/1 |
| decisions | 2/2 |
| stress | 4/4 |
| sleep | 1/1 |
| movement | 5/5 |
| memory | 2/2 |
| learning | 4/4 |
| thinking | 5/5 |
| communication | 4/4 |
| creativity | 3/3 |
| work | 3/3 |
| digital | 3/3 |
| multilingual | 3/3 |
| no-answer | 1/1 |
| safety | 2/2 |
| evidence-boundary | 2/2 |
| Case | Question | Gap |
|---|---|---|
| No current evaluation gaps. | ||
The usefulness score is a deterministic proxy. It does not grade writing quality, empathy, reasoning depth, or free-form hallucinations. Those belong in a separate model-level evaluation that consumes these same cases and expected boundaries.