Creative Evidence Lab
Can observable creative attributes be connected to performance evidence well enough to produce a change worth testing?
Nothing has been built or measured. No licensed creative assets and no aligned delivery logs are confirmed, and this page is a study design.
- Problem
- Creative intelligence products slide between three different claims: that a tag is accurate, that it predicts performance, and that changing the thing it describes improves the outcome.
- My contribution
- I designed the study that keeps them apart — an annotation protocol, a predictive evaluation with a fair baseline, and an experiment that intervenes on one attribute.
- Current result
- None. This is a design with stated stop conditions, published so the reasoning can be criticized before any of it is built.
The decision this has to serve
What would someone do differently because of it?
A creative strategist has to pick what the next batch of assets changes first. The useful output is not a score; it is one hypothesis with the evidence attached, the conditions it holds under, and the test that would kill it.
- 1 — Label quality
- Can two people apply the definition and agree? Until this holds, everything downstream is measuring annotation noise.
- 2 — Predictive value
- Does the label still add information out of sample, over a baseline that already knows the delivery context?
- 3 — Incremental impact
- Does changing the attribute move the outcome under a recorded random assignment? Nothing below this line supports the word "lift."
A pattern that clears the first two layers is still a hypothesis. The point of separating them is that they have different burdens of proof and are routinely reported as one.
One worked example
What does a single output look like?
Observed
The first spoken brand mention occurs at 00:06. The visual logo tag still requires review.
Hypothesis
An earlier brand reveal may improve recall for this campaign. Incremental conversion impact is not established.
Recommended test
Compare the current version against an earlier reveal, holding the offer and the CTA fixed.
Decision rule
Ship only if the preregistered primary outcome and the guardrails support the change; otherwise keep the original or collect more evidence.
The timestamps are illustrative. Every line in the product would be tagged as observation, association, hypothesis, or experimental result, so an uncertain suggestion cannot be rendered as a confident recommendation.
Design and data requirements
What gets built, in what order, and on what data?
One creative format first — a single language, fifteen to thirty second video. Not images, long video, landing pages, and cross-channel attribution in the same version.
| Tag group | Example fields | Annotation requirement |
|---|---|---|
| Attention | First product appearance, opening shot type, human voice in the first three seconds | Bound to a time range and a frame or transcript reference. |
| Branding | First visual logo time, first spoken brand time | Visual and spoken kept separate; absent and not-detected kept separate. |
| Connection | People on screen, demo or testimonial form, setting | Observable fields first. Anything like "warmth" needs a human definition and a reliability check. |
| Direction | CTA present, on-screen or spoken, specific action, first appearance | An explicit, observable action taxonomy. |
| Context | Duration, aspect ratio, language, channel, objective, creative family | Asset properties kept distinct from delivery context. |
Roughly eight to twelve core tags in a first version. Google’s ABCD framework is a public reference for organizing them, not ground truth and not a claim that each attribute has a general causal effect.
- Layer 1 — Is the label trustworthy?
- A rubric pilot of thirty to fifty assets, widening to roughly one hundred fifty to two hundred fifty, stratified by language, duration, and creative family. Part of the sample double-annotated, disagreements adjudicated. Report per-tag agreement and prevalence, never one blended score. Then compare OCR/ASR plus rules against a vision-language model on identical assets, with the final holdout isolated by creative family and frozen before testing.
- Layer 2 — Do the labels add predictive information?
- Only opens once assets and outcomes are genuinely aligned. Name one target — CTR, CVR, and ROAS are different tasks and do not roll up into a creative score. Baseline 0 is a grouped historical rate, Baseline 1 is pre-delivery context only, Model 2 adds tags. Out-of-time holdout, creative families isolated, calibration reported, and a leakage check for future performance leaking into tags or campaign names encoding the outcome.
- Layer 3 — Does changing the creative change the outcome?
- Two versions of the same asset differing only in the tested attribute, under a mechanism that records random assignment. Estimand: ITT difference in qualified conversion rate per assigned unit within a fixed window. Power computed from the real baseline and a business-defined minimum effect, with the stopping rule fixed in advance.
The data contract is five tables: creative_assets by asset version, creative_tags by asset × tag × model or human version, delivery_outcomes by asset × campaign × audience × placement × day, experiment_assignments by unit × experiment, and experiment_outcomes by unit over a fixed observation window.
Running two ad versions at the same time is not randomization — the platform may move budget and audience toward whichever is already winning. If only campaign- or geo-level randomization is available, the analysis and the power calculation have to use that cluster. And since exposure itself can be affected by treatment, click-conditioned CVR cannot stand in for the ITT result.
Current status
What exists today?
A design. No assets have been collected, no annotation has been run, and no performance data has been obtained.
| Stage | Condition to continue | If it fails |
|---|---|---|
| Taxonomy | Humans define and identify the core tags consistently. | Cut or rewrite the tags. Model consensus is not a substitute for reliability. |
| Tagging | Agreed accuracy and review cost met on the frozen holdout. | Keep a human or hybrid loop instead of chasing full automation. |
| Prediction | Aligned outcomes, a fair baseline, and out-of-sample gain. | Publish the negative result. Do not add performance-prediction language to it. |
| Improvement | An identifiable experiment runs and supports the change. | Ship testable hypotheses and claim no uplift. |
Route selection comes first: with usable assets and aligned logs this becomes a tagging-then-performance study; with assets only it becomes a tagging benchmark plus an annotation-efficiency trial, which is still a real result; with neither it is shelved rather than stalled on a data cold start.
Limitations
What this case does not establish.
- Nothing here has been built or measured. This page is a study design and says so at the top.
- Performance figures from my employment belong to that work and are not transferable to this project.
- Public ad libraries do not ship CTR, CVR, or ROAS alongside the creative. Stitching public creatives to unrelated click data would produce a join key that means nothing.