Writing
Notes on measurement and AI evaluation.
Practical notes on evaluating AI systems, measuring product outcomes, and learning from things I build. Each piece supports one claim, with the examples worked through rather than asserted.
When a Better Offline Metric Is the Wrong Launch Signal
A model can improve on the test set and still fail the launch decision. Two examples — recommendation quality and rare high-cost failures — show how the evaluation population, the metric, and the guardrails decide what an offline win actually means.
Creative Tagging Is Easy. Proving Lift Is Harder.
Finding a logo in a video, predicting which ad performs well, and proving that an earlier brand reveal improves outcomes are three different tasks with three different burdens of proof. This piece follows one creative change through all three.
An AI Evaluator Needs an Evaluation
An evaluator can reward a polished answer that misses the task. Two deliberately difficult answer pairs show how to test a judge for grounding and task completion, compare it against human ratings, and learn from the disagreements before its score gates a release.
In progress
What I am writing next
Listed with the claim each one has to support. They go up when the example work is done, not on a schedule.
- AI Evaluation
Too Many Narratives: Measuring Over-Merging and Over-Splitting
Two claims can mention the same company and still tell different stories. A small set of paired examples shows how to distinguish over-merging from over-splitting, write clearer annotation rules, and evaluate clusters without treating shared vocabulary as shared meaning.
- AI Evaluation
When Should an LLM Replace a Deterministic Rule?
A rule can be predictable and brittle; a model can be flexible and inconsistent. Using short claims with negation and stance changes, this piece sets out a comparison of rules, a model, and a hybrid — measuring errors, review effort, cost, and latency.
- Building
No Errors Is Not the Same as a Healthy Pipeline
Staleness, coverage against a required field set, and progress are the health metrics that catch silent failure. A pipeline with clean logs and no run records is failing quietly.
Research
Peer-reviewed publications
- 2022
Machine learning for food security: Principles for transparency and usability
Applied Economic Perspectives and Policy, 44 (2), 893–910
- 2020
Effects of stockholding policy on maize prices: Evidence from Zambia
Journal of Agricultural & Food Industrial Organization, 18 (1), 20190057
- 2019
A data-driven approach improves food insecurity crisis prediction
World Development, 122, 399–409