Yujun Zhou · Senior Data Scientist · Economics PhD

I measure and improve AI products.

I work on experimentation and evaluation for recommendations, advertising, and AI systems. Inside product teams I connect decisions to evidence; on my own I build the tools that test those ideas in practice.

Short-form video recommendations & integrity · Enterprise AI evaluation · Applied Economics PhD

Selected work

Decisions, and the evidence behind them

All work →

Writing

Notes on measurement and AI evaluation

All writing →
  • When a Better Offline Metric Is the Wrong Launch Signal

    A model can improve on the test set and still fail the launch decision. Two examples — recommendation quality and rare high-cost failures — show how the evaluation population, the metric, and the guardrails decide what an offline win actually means.

  • Creative Tagging Is Easy. Proving Lift Is Harder.

    Finding a logo in a video, predicting which ad performs well, and proving that an earlier brand reveal improves outcomes are three different tasks with three different burdens of proof. This piece follows one creative change through all three.

  • An AI Evaluator Needs an Evaluation

    An evaluator can reward a polished answer that misses the task. Two deliberately difficult answer pairs show how to test a judge for grounding and task completion, compare it against human ratings, and learn from the disagreements before its score gates a release.

Contact

Let’s talk about product measurement and AI evaluation.

I’m interested in teams turning experimentation and evaluation into better product decisions.