← All work

Study designIndependent build

Creative Evidence Lab

Can observable creative attributes be connected to performance evidence well enough to produce a change worth testing?

Nothing has been built or measured. No licensed creative assets and no aligned delivery logs are confirmed, and this page is a study design.

Problem
Creative intelligence products slide between three different claims: that a tag is accurate, that it predicts performance, and that changing the thing it describes improves the outcome.
My contribution
I designed the study that keeps them apart — an annotation protocol, a predictive evaluation with a fair baseline, and an experiment that intervenes on one attribute.
Current result
None. This is a design with stated stop conditions, published so the reasoning can be criticized before any of it is built.
My role
Proposed sole author: taxonomy, annotation protocol, tagging baselines, error analysis, and experiment design.
Methods
Multimodal tagging · Annotator agreement · Predictive evaluation · Experiment design
Evidence
None yet — this is a study design.
01

The decision this has to serve

What would someone do differently because of it?

A creative strategist has to pick what the next batch of assets changes first. The useful output is not a score; it is one hypothesis with the evidence attached, the conditions it holds under, and the test that would kill it.

1 — Label quality
Can two people apply the definition and agree? Until this holds, everything downstream is measuring annotation noise.
2 — Predictive value
Does the label still add information out of sample, over a baseline that already knows the delivery context?
3 — Incremental impact
Does changing the attribute move the outcome under a recorded random assignment? Nothing below this line supports the word "lift."

A pattern that clears the first two layers is still a hypothesis. The point of separating them is that they have different burdens of proof and are routinely reported as one.

02

One worked example

What does a single output look like?

Observed

The first spoken brand mention occurs at 00:06. The visual logo tag still requires review.

Hypothesis

An earlier brand reveal may improve recall for this campaign. Incremental conversion impact is not established.

Recommended test

Compare the current version against an earlier reveal, holding the offer and the CTA fixed.

Decision rule

Ship only if the preregistered primary outcome and the guardrails support the change; otherwise keep the original or collect more evidence.

The timestamps are illustrative. Every line in the product would be tagged as observation, association, hypothesis, or experimental result, so an uncertain suggestion cannot be rendered as a confident recommendation.

03

Design and data requirements

What gets built, in what order, and on what data?

One creative format first — a single language, fifteen to thirty second video. Not images, long video, landing pages, and cross-channel attribution in the same version.

Tag groupExample fieldsAnnotation requirement
AttentionFirst product appearance, opening shot type, human voice in the first three secondsBound to a time range and a frame or transcript reference.
BrandingFirst visual logo time, first spoken brand timeVisual and spoken kept separate; absent and not-detected kept separate.
ConnectionPeople on screen, demo or testimonial form, settingObservable fields first. Anything like "warmth" needs a human definition and a reliability check.
DirectionCTA present, on-screen or spoken, specific action, first appearanceAn explicit, observable action taxonomy.
ContextDuration, aspect ratio, language, channel, objective, creative familyAsset properties kept distinct from delivery context.

Roughly eight to twelve core tags in a first version. Google’s ABCD framework is a public reference for organizing them, not ground truth and not a claim that each attribute has a general causal effect.

Layer 1 — Is the label trustworthy?
A rubric pilot of thirty to fifty assets, widening to roughly one hundred fifty to two hundred fifty, stratified by language, duration, and creative family. Part of the sample double-annotated, disagreements adjudicated. Report per-tag agreement and prevalence, never one blended score. Then compare OCR/ASR plus rules against a vision-language model on identical assets, with the final holdout isolated by creative family and frozen before testing.
Layer 2 — Do the labels add predictive information?
Only opens once assets and outcomes are genuinely aligned. Name one target — CTR, CVR, and ROAS are different tasks and do not roll up into a creative score. Baseline 0 is a grouped historical rate, Baseline 1 is pre-delivery context only, Model 2 adds tags. Out-of-time holdout, creative families isolated, calibration reported, and a leakage check for future performance leaking into tags or campaign names encoding the outcome.
Layer 3 — Does changing the creative change the outcome?
Two versions of the same asset differing only in the tested attribute, under a mechanism that records random assignment. Estimand: ITT difference in qualified conversion rate per assigned unit within a fixed window. Power computed from the real baseline and a business-defined minimum effect, with the stopping rule fixed in advance.

The data contract is five tables: creative_assets by asset version, creative_tags by asset × tag × model or human version, delivery_outcomes by asset × campaign × audience × placement × day, experiment_assignments by unit × experiment, and experiment_outcomes by unit over a fixed observation window.

Running two ad versions at the same time is not randomization — the platform may move budget and audience toward whichever is already winning. If only campaign- or geo-level randomization is available, the analysis and the power calculation have to use that cluster. And since exposure itself can be affected by treatment, click-conditioned CVR cannot stand in for the ITT result.

04

Current status

What exists today?

A design. No assets have been collected, no annotation has been run, and no performance data has been obtained.

StageCondition to continueIf it fails
TaxonomyHumans define and identify the core tags consistently.Cut or rewrite the tags. Model consensus is not a substitute for reliability.
TaggingAgreed accuracy and review cost met on the frozen holdout.Keep a human or hybrid loop instead of chasing full automation.
PredictionAligned outcomes, a fair baseline, and out-of-sample gain.Publish the negative result. Do not add performance-prediction language to it.
ImprovementAn identifiable experiment runs and supports the change.Ship testable hypotheses and claim no uplift.

Route selection comes first: with usable assets and aligned logs this becomes a tagging-then-performance study; with assets only it becomes a tagging benchmark plus an annotation-efficiency trial, which is still a real result; with neither it is shelved rather than stalled on a data cold start.

05

Limitations

What this case does not establish.

  • Nothing here has been built or measured. This page is a study design and says so at the top.
  • Performance figures from my employment belong to that work and are not transferable to this project.
  • Public ad libraries do not ship CTR, CVR, or ROAS alongside the creative. Stitching public creatives to unrelated click data would produce a join key that means nothing.