Making Enterprise AI Evaluation Launch-Relevant
How do you stop an assistant that sounds right from failing the task it was built for?
Internal enterprise work. No client data, prompts, or transcripts are published here; the failure trace below is a reconstruction written for this page.
- Problem
- Answer-quality review kept passing systems that still could not finish the task a strategist actually had. Reviewers were grading prose; the product had to deliver a decision.
- My contribution
- I defined the evaluation as four separate layers — retrieval, tool use, unsupported claims, end-to-end task success — built the test set and rubric behind them, and set the release gate.
- What changed
- Release stopped depending on demo impressions and started depending on a frozen test set with a threshold stated before the run.
The problem
Why did a good review score not predict a working product?
The failure mode that started this: a system that reviewed well answer by answer, and still could not complete a task end to end. A reviewer reading one answer at a time is grading fluency, sourcing, and tone. None of those tell you whether the person who asked could then go and do the thing they came to do.
A single answer-quality score also averages across failures that have nothing in common. Retrieval missing a document and a model misreading a document it did retrieve are the same score and completely different repairs.
My contribution
What was measured, layer by layer?
| Layer | What it asks | What a good score does not prove |
|---|---|---|
| Retrieval | Did the right source reach the model at all? | That the model then used it. |
| Tool use | Was the right tool called with the right arguments? | That the returned value was interpreted correctly. |
| Unsupported claims | Is every factual sentence traceable to retrieved evidence? | That the supported sentences answer the question. |
| End-to-end success | Could the user finish the task without repair? | That it generalizes past the test set. |
Each layer has its own denominator. Unsupported-claim rate is per factual sentence; task success is per task. A system can improve sharply on one while the other is flat, and a single blended number hides exactly the tradeoff a release decision turns on.
Alongside the layers I built the test set from real task families rather than prompts that happened to demo well, wrote the rubric that defines each score point, and set the release gate as a threshold on a frozen slice, stated before the run together with what meeting it costs in latency and spend.
One failure, traced end to end
What does a layered diagnosis actually look like?
Illustrative reconstruction. This trace is written for this page using a self-authored task. It shows the shape of the diagnosis, not a real client incident.
Task
Which of our three campaigns should absorb the next increment of budget, and on what evidence?
Retrieval
Correct. The performance summary for all three campaigns was retrieved and passed to the model.
Tool use
Correct. The spend query ran with the right date range and returned the right rows.
Answer
Fluent, cited, and wrong at the level that matters: it quoted an adjacent passage describing a different campaign period, and recommended the increment on that basis.
Located failure
Not a model-capability failure. The chunk boundary split the campaign label from the numbers beneath it, so a correctly retrieved document supported a confident sentence about the wrong thing.
Repair
Chunking that keeps a label with its table, and a citation constraint that requires the cited span to contain the entity being described.
A trace like this changes how a team argues. "The model is bad at this" becomes a claim you have to locate in a layer — and in my experience most located failures are not model failures.
Method
What keeps the evaluation itself trustworthy?
- The test set is stratified by task family so no single client or task type dominates the score.
- The rubric defines each score point with worked examples at the boundaries. Where reviewers cannot agree, the rubric gets rewritten or the category dropped, rather than averaged over.
- A held-out slice stays frozen. Prompt iteration happens on the development slice, or the gate is measuring the tuning rather than the system.
- An LLM judge is validated against human labels before its score is allowed to gate anything. Agreement with itself is not correctness, and a judge sharing the generator’s blind spots will certify them.
- Quality is quoted with its cost and latency. A configuration that wins on quality and doubles response time is a tradeoff for the product owner, not a win to announce.
The tension worth stating plainly: tightening an evidence requirement raises factuality and also raises refusals, including refusals on tasks the system could have completed. Coverage and factuality trade against each other, so the gate has to be derived from the cost of each error type rather than set at whatever number looks rigorous.
What a technical reader can ask me
Where does the detail live?
- The rubric structure, and how boundary cases get adjudicated.
- How to validate an LLM judge against human labels, and what to do where they keep disagreeing.
- How to derive a regression threshold from the cost of each error type instead of picking a round number.
Limitations
What this case does not establish.
- The layers have different tasks and different denominators. They do not combine into one accuracy number, and I do not report one.
- A frozen test set ages. It measures the failures we already knew to look for, which is why the taxonomy feeds new cases back into it.
- Human rubric scores carry annotator disagreement. A mean without an agreement rate overstates precision.
- The end-to-end trace in this case is a reconstruction built to show the method. It is not a client incident report.
- Reported experience
- Described from work I did inside a company. The underlying data is not public and is not reproduced here.