Creative Tagging Is Easy. Proving Lift Is Harder.
Finding a logo in a video, predicting which ad performs well, and proving that an earlier brand reveal improves outcomes are three different tasks with three different burdens of proof. This piece follows one creative change through all three.
Take a fifteen-second video ad. The logo appears at 00:02 in the corner of the frame. A voiceover says the brand name at 00:06. A call to action appears on screen at 00:12 and is spoken at 00:13.
A creative intelligence product will tell you three things about this ad, and they sound like one thing. They are not. Accurate tags, predictive tags, and tags whose underlying attribute causes the outcome are separate claims, and each needs its own evaluation.
Claim one: the tag is correct
Ask two annotators for "first brand appearance" on the ad above. One says 00:02, reading the corner logo. The other says 00:06, reasoning that a logo small enough to miss is not an appearance in any sense the viewer experiences. Both are being careful. The definition is what failed.
This is the cheapest failure to fix and the one most often skipped, because the model produces a number either way. The repairs are unglamorous: split visual and spoken brand presence into separate fields; require every tag to be bound to a time range and a frame or transcript reference; keep "absent" and "not detected" as different values, because a pipeline that cannot see the logo and an ad that has no logo are different facts about the world.
Then measure per tag, not in aggregate. A blended tagging accuracy of 91% is compatible with a CTA-presence tag at 99% and a first-brand-appearance tag at 62%, and the second one is the tag every downstream analysis is about to use.
Zero errors on a pilot of thirty assets is not a zero error rate. Report the sample size and the interval next to every accuracy figure, or the number is decoration.
Claim two: the tag predicts performance
Suppose the labels are now trustworthy, and the data shows that ads with a brand reveal before 00:03 have higher ROAS than ads that reveal later. This is a real pattern in the data. It is also close to uninterpretable on its own.
Large brands with big budgets reveal early — it is house style, and they have the recognition to make it work. They also buy better placements, target warmer audiences, and run against higher-converting offers. Every one of those is a plausible explanation for the ROAS gap, and the tag is partly a proxy for "this is a large advertiser."
A predictive evaluation that means anything has to answer: better than what? The honest baseline ladder is three rungs.
| Model | What it knows | What it establishes |
|---|---|---|
| Baseline 0 | The grouped historical rate for this advertiser and objective | The floor. A tag model that loses to this has added nothing. |
| Baseline 1 | Pre-delivery context: budget, audience, placement, objective, format | How much of the pattern is delivery context wearing a creative costume. |
| Model 2 | Baseline 1 plus the creative tags | The incremental information in the tags, and only that. |
Hold out by creative family and out of time, or near-duplicate assets from the same campaign land on both sides of the split and the score measures memorization. Check for leakage in both directions: performance data that leaked into a tag, and campaign names that encode the outcome. And pick one target — CTR, CVR, and ROAS are different tasks with different denominators, and they do not roll up into a creative score.
Suppose Model 2 beats Baseline 1 out of sample. You have shown the tags carry information the delivery context does not. You have not shown that changing the creative changes anything.
Claim three: changing the attribute changes the outcome
Only an intervention gets you here, and the intervention has to be narrow enough to interpret. Two versions of the same asset, differing in the tested attribute and nothing else.
Observation
In this account, ads with a spoken brand mention before 00:03 show higher ROAS than ads that mention it later. Confounded with budget, audience, and offer.
Hypothesis
Moving the spoken brand mention earlier improves conversion rate for this campaign. Not established.
Intervention
Version A: brand mention at 00:06, unchanged. Version B: identical edit with the brand mention at 00:02. Same offer, same CTA, same length, same audio bed.
Metric
ITT difference in qualified conversion rate per assigned unit, in a fixed window. Guardrails on completion rate and cost per qualified action.
Decision rule
Ship B only if the preregistered primary outcome clears the minimum effect and no guardrail is crossed. Otherwise keep A.
Two traps at this stage. Running both versions simultaneously in one campaign is not randomization — the delivery system reallocates budget toward whichever version is already winning, so exposure is a function of the outcome. If the only randomization available is campaign or geo level, then that cluster is the unit, and the power calculation has to use it. Second, click-conditioned conversion rate cannot stand in for the ITT result, because treatment can change who clicks in the first place.
The check
Before repeating any creative finding, ask which of the three claims it actually is:
- Is this a label, and do I know its per-tag agreement and sample size?
- Is this a prediction, and does it beat a baseline that already knows the delivery context?
- Is this a causal claim, and is there a recorded random assignment behind it?
Most creative insight lives at rung one or two and gets described in the language of rung three. Accurate tags do not establish incremental lift. Each claim needs its own evaluation, and saying which one you have is not a hedge — it is the entire content of the finding.