The Metric That Could Not See It
A grounding metric scored a hallucinated answer a perfect 1.00 — and the bug was not in the metric.
Adding automated grounding metrics to the evo-ai evaluation loop came with one condition: before any of them was allowed to fail a build, each had to prove it could see a defect we already knew about. One of them could not — and chasing that down found a real bug in the product, in the one place nobody looks: the evidence an answer is graded against.
The bar, set before the code
Grounding metrics are seductive. A library — here, DeepEval — scores an answer for faithfulness, hands you a number between zero and one, and it is extremely tempting to put that number in a pipeline and call the problem solved.
So the plan for the stage carried a condition, written down before any of it was built: the metrics must be able to see the grounding defects we had already found, on the original answers, and clear the fixed ones. If they cannot see those, they are not worth enforcing.
Two defects were available, both already diagnosed, fixed, and written up at the time:
- A hallucinated count. Asked how many permits the workspace held, the assistant had answered "There are 2 permits." There were eighty-five.
- Zero rows reported as a blanket refusal. Asked for current plant temperatures, it had answered "This workspace's records don't cover that question…" — when the records covered it perfectly well.
Both are the failure mode that matters most in a regulated setting. Neither looks like an error. Both are delivered in the same calm tone as a correct answer.
Replaying a defect that was already fixed
The problem with testing a detector against a fixed bug is that the bug is fixed. You cannot ask the live service for the wrong answer any more.
The harness — evo.ragframework, built for this loop — has a replay adapter for exactly this: instead of calling a service, it feeds recorded answers and recorded sources through the identical scoring path. So the old answers — quoted verbatim from the write-ups, since the original run's data was long gone — were replayed against the fixed answers' sources. Same chunks, same evidence, same judge, same code path. The only thing that differed between the two runs was the text of the answer.
That is the whole trick, and it is worth stealing: if you want to know whether a check works, give it a known-bad input and a known-good one that differ in exactly one way.
A perfect score for a wrong answer
Gemini 2.5 Pro as the metric model, three trials each, median kept:
| Answer | Faithfulness | Verdict |
|---|---|---|
| "There are 2 permits." (the hallucination) | 1.00 | FAIL (by the judge) |
| "There are 85 permits in total…" (the fix) | 1.00 | PASS |
| "…records don't cover that question" (the false refusal) | 0.00 | FAIL |
| "Here are the current temperatures…" (the fix) | 1.00 | PASS |
The hallucinated count scored a perfect 1.00. Faithfulness measures whether an answer's claims are supported by the context it was given, and it found no contradiction at all between "there are 2 permits" and the evidence — because the evidence it was handed contained no count whatsoever.
Had that stage been wired into a pipeline on the strength of "we added grounding metrics", it would have been a gate that could never have caught the defect it was bought for. It would have been worse than nothing, because it would have been reassuring.
Grading against the instructions
The assistant answers a counting question through a SQL route: a planner writes a query, the query runs against curated read-only views, and the result becomes the context the answer is written from. When the API returns that answer it also returns its sources, so a caller can show the evidence.
For an analytics answer, that source text was the first five hundred characters of the result node — and the result node opened with the planner's own instruction preamble. Five hundred characters of guidance about how to phrase a result, and then the text was cut off. The rows and the aggregate never appeared in it.
So the metric was doing its job correctly on the wrong input. It was asked whether "there are 2 permits" is supported by a block of prompt instructions, and the honest answer to that is that the instructions neither support nor contradict it. 1.00.
This is the part I would not have predicted: the defect was not in the library, the model, the threshold or the harness. It was in what the product called evidence — a field nobody reads closely, because it is usually rendered as a citation chip and glanced at.
The other result was the real tell
The false refusal scored 0.00 — a correct verdict. It would have been easy to take that as partial success and move on.
But reading the reason rather than the number showed it was right by accident. The metric was not contradicting the refusal against the plant temperatures; it was contradicting it against a line in the preamble that told the model never to say "don't cover". The answer disagreed with its own instructions, and the metric noticed that, which is not the thing anyone wanted measured.
A correct verdict for an incorrect reason is a failing test that happens to be green. It will pass until the prompt is reworded, and then it will silently stop working. That reason line was worth more than either score.
Fixing the product, not the metric
The temptation here is to fix the measurement: feed the metric a different field, widen the character limit, special-case the analytics route. All of that would have made the number go the right way and left the product wrong.
Because this was not only an evaluation problem. The same truncated text is what the API hands any caller that wants to show its working — so the product's own evidence for a numeric answer was its instructions rather than its data. It was filed as a product bug and fixed there: for an analytics answer, the source text now opens with the query, the row count and the rows.
Then the identical replay ran again, against the new evidence:
| Answer | Faithfulness before | after |
|---|---|---|
| "There are 2 permits." | 1.00 | 0.00 (0, 0, 0) |
| "There are 85 permits in total…" | 1.00 | 1.00 (1, 1, 1) |
| "…records don't cover that question" | 0.00 (wrong reason) | 0.00 (0, 0, 0) |
| "Here are the current temperatures…" | 1.00 | 1.00 (1, 1, 1) |
And the reasons were finally about the records. On the hallucinated count:
the output incorrectly states there are only two permits, while the context explicitly mentions 85 permits
That is the sentence the whole exercise was for. Not the number — the number was available the first time and was wrong. The sentence.
Three other things that fell out
- Answer relevancy is blind to both defects, and should never be gated. It scored 1.00 everywhere — the hallucination, the refusal, the fixes. A confidently wrong number is entirely relevant to the question it answers. In one trial it scored a correct answer 0.2, for including "irrelevant details about specific permits" — the permits the question had asked it to name.
- One trial is not a measurement. Three trials of one metric on one unchanged answer came back 1.0, 1.0 and 0.2; another gave 0.5, 1.0, 1.0. The median was stable and any single trial was a coin flip. Everything recorded is a median of three, and a stage that ran once per answer would have produced a changelog of phantom regressions.
- Read the metric's definition, not its name. DeepEval’s hallucination metric scores the share of the context the answer agrees with — so 1.00 means nothing was contradicted, the opposite of what the name suggests to anyone skim-reading a dashboard. The stage now records each metric's own verdict rather than comparing a raw score against a threshold, specifically so a mis-assumed direction cannot become a gate.
What I would tell the next person
A metric that returns a number is not a metric that works. The only way to know is to show it a defect you already understand and watch what it does. That costs an afternoon and it is the difference between a quality gate and a reassurance mechanism.
Read the reason, not the score. Both scores in the first run were defensible. One was right for the wrong reason and one was wrong for a good reason, and only the written explanations distinguished them. A dashboard of numbers with the reasons collapsed would have hidden the entire finding.
When a measurement disagrees with you, the input is the first suspect. Not the library, not the model, not the threshold. The metric here was correct throughout; it was being handed the wrong text, by the product, in a field that was wrong for every consumer and not just for the test.
And the thing I keep relearning: building the measurement found a bug that the thing being measured had been shipping for months. Instrumenting a system honestly is itself a way of testing it.
The stack, for the record
Harness: evo.ragframework, built for this loop, with httpx for the HTTP adapter, SQLite for every run, and pytest on GitHub Actions for its own suite. Judging and metrics: DeepEval 4.2 — Faithfulness, Answer Relevancy, Contextual Precision and Recall, Hallucination and G-Eval — with Gemini 2.5 Pro as the judge, reached through LiteLLM, and Ollama running Llama 3.1 8B for offline plumbing. Tracing: Langfuse 4.50, self-hosted on ClickHouse, PostgreSQL, Redis and MinIO, fed by OpenTelemetry and OpenInference from inside the service. Runtime: Docker Compose on Docker Desktop, and uv, so the service is built from exactly the dependency set that was tested. Under test: evo-ai — FastAPI, LlamaIndex, Qdrant with fastembed, and LiteLLM or Ollama for generation.
What each one is for is set out in the tooling section of the case study.