← All work

Study designIndependent build

Analytics Skeptic

Can an AI reviewer find the metric problems that would change a launch decision, rather than just sounding reasonable?

A prompt and a set of war stories. Not deployed, and nothing has been benchmarked.

Problem
An AI reviewer that sounds reasonable is easy. One that catches the specific issue that would change a launch decision is the only version worth having.
My contribution
I wrote the prompt and the war stories behind it, and designed the blind evaluation that would separate useful criticism from confident noise.
Current result
A prompt, not a product. The honest next step is five real users, not a bigger architecture.
My role
Proposed: benchmark construction, blind scoring design, and comparison against a human checklist.
Methods
Blind evaluation · Benchmark construction · Adjudicated labels
Evidence
None yet — this is a study design.
01

The decision this has to serve

What exists, and what is the open question?

A prompt and a collection of war stories about analyses that went wrong. The README states the position plainly: no deployment, and the goal is five real users first. That is where it stands.

The question is not whether the output reads well. It is whether it surfaces the issue that would have changed the decision — and whether it does that more often than a fixed human checklist.

02

What a real evaluation would look like

How would it be tested?

  • Thirty to fifty publishable cases covering wrong OEC, drifting denominators, selection bias, guardrail conflicts, and insufficient information.
  • Isolate development and test by scenario family, so a variant of a known case cannot leak across the split.
  • Compare three arms: a fixed human checklist, a generic prompt, and the skeptic prompt.
  • Score blind on recall of the decision-changing issue, rate of unfounded criticism, evidence faithfulness, actionability, and whether the correct decision changed.
  • Bootstrap by case, and keep the author of the gold standard out of the scoring.
03

Next

What happens next?

Finish the five-user trial. If people reuse it and it beats a plain checklist, build the frozen benchmark and the comparison report. If they do not, it stays a prompt — which is a perfectly good outcome for a prompt.

04

Limitations

What this case does not establish.

  • Nothing is deployed. Describing this as an agent platform would be false.
  • A single author cannot write the gold standard and also score against it. Blind scoring needs someone else.
  • "Users found it insightful" is a separate record from whether it was correct, and cannot substitute for it.