Writing

Notes on measurement and AI evaluation.

Practical notes on evaluating AI systems, measuring product outcomes, and learning from things I build. Each piece supports one claim, with the examples worked through rather than asserted.

In progress

What I am writing next

Listed with the claim each one has to support. They go up when the example work is done, not on a schedule.

  • AI Evaluation

    Too Many Narratives: Measuring Over-Merging and Over-Splitting

    Two claims can mention the same company and still tell different stories. A small set of paired examples shows how to distinguish over-merging from over-splitting, write clearer annotation rules, and evaluate clusters without treating shared vocabulary as shared meaning.

  • AI Evaluation

    When Should an LLM Replace a Deterministic Rule?

    A rule can be predictable and brittle; a model can be flexible and inconsistent. Using short claims with negation and stance changes, this piece sets out a comparison of rules, a model, and a hybrid — measuring errors, review effort, cost, and latency.

  • Building

    No Errors Is Not the Same as a Healthy Pipeline

    Staleness, coverage against a required field set, and progress are the health metrics that catch silent failure. A pipeline with clean logs and no run records is failing quietly.

Research

Peer-reviewed publications

Google Scholar →