Recommendation Quality Beyond Engagement
Defined consumption-side measures of viewer experience and made them part of how the team read a launch.
Yujun Zhou · Senior Data Scientist · Economics PhD
I work on experimentation and evaluation for recommendations, advertising, and AI systems. Inside product teams I connect decisions to evidence; on my own I build the tools that test those ideas in practice.
Short-form video recommendations & integrity · Enterprise AI evaluation · Applied Economics PhD
Selected work
Defined consumption-side measures of viewer experience and made them part of how the team read a launch.
Split one answer-quality score into layers that fail differently, then tied release to a stated gate instead of a demo.
Built a pipeline that connects changing claims to their sources and tracks how they evolve over time.
Writing
A model can improve on the test set and still fail the launch decision. Two examples — recommendation quality and rare high-cost failures — show how the evaluation population, the metric, and the guardrails decide what an offline win actually means.
Finding a logo in a video, predicting which ad performs well, and proving that an earlier brand reveal improves outcomes are three different tasks with three different burdens of proof. This piece follows one creative change through all three.
An evaluator can reward a polished answer that misses the task. Two deliberately difficult answer pairs show how to test a judge for grounding and task completion, compare it against human ratings, and learn from the disagreements before its score gates a release.
Creative Evidence Lab · Study design
A proposed workflow connecting observable creative attributes, performance hypotheses, and an experiment brief — with label quality, predictive value, and incremental impact kept as three separate claims.
Contact
I’m interested in teams turning experimentation and evaluation into better product decisions.