When a Better Offline Metric Is the Wrong Launch Signal
A model can improve on the test set and still fail the launch decision. Two examples — recommendation quality and rare high-cost failures — show how the evaluation population, the metric, and the guardrails decide what an offline win actually means.
Here is a readout that arrives on a launch review most weeks. The new model scores 83% against 80% for the incumbent on the held-out set. It is a clean win, the sample is large, and everyone in the room would like to ship it.
Then someone splits the same evaluation set into the ordinary slice and the slice the product actually worries about.
| Slice | Share of population | Incumbent | Candidate |
|---|---|---|---|
| Ordinary traffic | 95% | 80% | 83% |
| High-risk slice | 5% | 60% | 50% |
| Weighted overall | 100% | 79.0% | 81.35% |
Illustrative numbers, chosen so the arithmetic is easy to check. The slice weights are the population shares, not a modelling choice.
The candidate wins overall by more than two points and loses the high-risk slice by ten. Both statements are true, computed from the same table. Which one is the launch signal depends on a question the accuracy number cannot answer: what is this evaluation supposed to decide?
Overall performance and slice performance answer different questions
A sample drawn from typical traffic is the right sample for estimating typical performance. That is not a flaw, it is the design working. What it cannot do is say much about a slice that makes up 5% of the population — at that share, most of the sample is spent on the case you already understand, and the interval on the slice estimate is wide enough to hide a real regression.
The mistake is not sampling from typical traffic. The mistake is reading one number as though it answered both questions. If the decision depends on the rare slice, then the slice needs enough sample to support an estimate, which usually means oversampling it and weighting back when you want the population figure. Reporting the weighted overall alongside the unweighted slice estimates, each with its interval, costs one extra row and removes the ambiguity entirely.
A point estimate without a sample size is not a result. Two of the four numbers in the table above would move meaningfully on a resample of the 5% slice, and nothing in this article tells you which two.
The metric can be gameable in a direction the test set cannot see
The second failure is subtler, because it survives every sampling fix. Consider a recommendation model evaluated on whether the viewer engages with what it shows. A model that leans harder on the creators a viewer has already reacted to will score well on that metric. It is genuinely better at predicting engagement.
It may also be showing the same three creators repeatedly. Engagement per item goes up. The session gets narrower. Offline, those two outcomes are indistinguishable, because the offline metric is built from historical interactions with a catalogue the model is now choosing differently from.
The fix is not a better offline metric. It is knowing which product failure the offline metric is structurally unable to see, and putting that failure on the launch readout as a separate measurement — repetition, supply concentration, the effort a viewer spends before finding something worth watching. These are proxies for experience, and imperfect ones. Their value is that they fail in a different direction from the primary metric, so when both move the right way you have learned something the primary metric alone could not tell you.
A guardrail needs a threshold and a consequence
It is tempting to conclude that a guardrail which never stops a launch is decoration. That inference does not hold. A guardrail may never bind because the risk did not materialize, or because the design that produced the candidate already avoided it. Neither is evidence the guardrail is useless.
What does make a guardrail decoration is being unfalsifiable: no defined failure mode, no stated threshold, and no consequence attached to crossing it. Those three are checkable before the launch, and they are what separates a guardrail from a chart. A guardrail also needs enough sensitivity to detect harm at a size that would actually change the decision — a metric too noisy to move outside its own interval will never bind no matter what happens.
What a launch readout has to state
The version of this I would defend in a review is short. Four fields, stated before the results are read:
- Population — who this was evaluated on, and how that differs from who the decision affects. Name the slices before you look.
- Comparison — what the candidate is being compared against, and whether the comparison holds anything constant that the launch will not.
- Primary metric — one, with its denominator, plus what it structurally cannot see.
- Guardrails — each with a failure mode, a threshold, and what crossing it triggers. "Escalate for discussion" is a legitimate consequence; "note it" is not.
None of this makes an offline win worthless. It makes the win a claim about a specific population and a specific outcome, which is the only form in which it can be checked. The launch decision is a different question, and it stays a different question no matter how clean the test-set number is.