← All work

Case studyAdvertising and media group

Making Enterprise AI Evaluation Launch-Relevant

How do you stop an assistant that sounds right from failing the task it was built for?

Internal enterprise work. No client data, prompts, or transcripts are published here; the failure trace below is a reconstruction written for this page.

Problem
Answer-quality review kept passing systems that still could not finish the task a strategist actually had. Reviewers were grading prose; the product had to deliver a decision.
My contribution
I defined the evaluation as four separate layers — retrieval, tool use, unsupported claims, end-to-end task success — built the test set and rubric behind them, and set the release gate.
What changed
Release stopped depending on demo impressions and started depending on a frozen test set with a threshold stated before the run.
My role
Defined the evaluation, built the test set and rubric, set the release gate, and drove adoption across the team.
Methods
Layered task evaluation · Human rubrics · Failure taxonomy · Regression gates · Cost/latency tradeoffs
Evidence
Reported experience
01

The problem

Why did a good review score not predict a working product?

The failure mode that started this: a system that reviewed well answer by answer, and still could not complete a task end to end. A reviewer reading one answer at a time is grading fluency, sourcing, and tone. None of those tell you whether the person who asked could then go and do the thing they came to do.

A single answer-quality score also averages across failures that have nothing in common. Retrieval missing a document and a model misreading a document it did retrieve are the same score and completely different repairs.

02

My contribution

What was measured, layer by layer?

LayerWhat it asksWhat a good score does not prove
RetrievalDid the right source reach the model at all?That the model then used it.
Tool useWas the right tool called with the right arguments?That the returned value was interpreted correctly.
Unsupported claimsIs every factual sentence traceable to retrieved evidence?That the supported sentences answer the question.
End-to-end successCould the user finish the task without repair?That it generalizes past the test set.

Each layer has its own denominator. Unsupported-claim rate is per factual sentence; task success is per task. A system can improve sharply on one while the other is flat, and a single blended number hides exactly the tradeoff a release decision turns on.

Alongside the layers I built the test set from real task families rather than prompts that happened to demo well, wrote the rubric that defines each score point, and set the release gate as a threshold on a frozen slice, stated before the run together with what meeting it costs in latency and spend.

03

One failure, traced end to end

What does a layered diagnosis actually look like?

Illustrative reconstruction. This trace is written for this page using a self-authored task. It shows the shape of the diagnosis, not a real client incident.

Task

Which of our three campaigns should absorb the next increment of budget, and on what evidence?

Retrieval

Correct. The performance summary for all three campaigns was retrieved and passed to the model.

Tool use

Correct. The spend query ran with the right date range and returned the right rows.

Answer

Fluent, cited, and wrong at the level that matters: it quoted an adjacent passage describing a different campaign period, and recommended the increment on that basis.

Located failure

Not a model-capability failure. The chunk boundary split the campaign label from the numbers beneath it, so a correctly retrieved document supported a confident sentence about the wrong thing.

Repair

Chunking that keeps a label with its table, and a citation constraint that requires the cited span to contain the entity being described.

A trace like this changes how a team argues. "The model is bad at this" becomes a claim you have to locate in a layer — and in my experience most located failures are not model failures.

04

Method

What keeps the evaluation itself trustworthy?

  • The test set is stratified by task family so no single client or task type dominates the score.
  • The rubric defines each score point with worked examples at the boundaries. Where reviewers cannot agree, the rubric gets rewritten or the category dropped, rather than averaged over.
  • A held-out slice stays frozen. Prompt iteration happens on the development slice, or the gate is measuring the tuning rather than the system.
  • An LLM judge is validated against human labels before its score is allowed to gate anything. Agreement with itself is not correctness, and a judge sharing the generator’s blind spots will certify them.
  • Quality is quoted with its cost and latency. A configuration that wins on quality and doubles response time is a tradeoff for the product owner, not a win to announce.

The tension worth stating plainly: tightening an evidence requirement raises factuality and also raises refusals, including refusals on tasks the system could have completed. Coverage and factuality trade against each other, so the gate has to be derived from the cost of each error type rather than set at whatever number looks rigorous.

05

What a technical reader can ask me

Where does the detail live?

  • The rubric structure, and how boundary cases get adjudicated.
  • How to validate an LLM judge against human labels, and what to do where they keep disagreeing.
  • How to derive a regression threshold from the cost of each error type instead of picking a round number.
06

Limitations

What this case does not establish.

  • The layers have different tasks and different denominators. They do not combine into one accuracy number, and I do not report one.
  • A frozen test set ages. It measures the failures we already knew to look for, which is why the taxonomy feeds new cases back into it.
  • Human rubric scores carry annotator disagreement. A mean without an agreement rate overstates precision.
  • The end-to-end trace in this case is a reconstruction built to show the method. It is not a client incident report.
Reported experience
Described from work I did inside a company. The underlying data is not public and is not reproduced here.