Analytics Skeptic
Can an AI reviewer find the metric problems that would change a launch decision, rather than just sounding reasonable?
A prompt and a set of war stories. Not deployed, and nothing has been benchmarked.
- Problem
- An AI reviewer that sounds reasonable is easy. One that catches the specific issue that would change a launch decision is the only version worth having.
- My contribution
- I wrote the prompt and the war stories behind it, and designed the blind evaluation that would separate useful criticism from confident noise.
- Current result
- A prompt, not a product. The honest next step is five real users, not a bigger architecture.
The decision this has to serve
What exists, and what is the open question?
A prompt and a collection of war stories about analyses that went wrong. The README states the position plainly: no deployment, and the goal is five real users first. That is where it stands.
The question is not whether the output reads well. It is whether it surfaces the issue that would have changed the decision — and whether it does that more often than a fixed human checklist.
What a real evaluation would look like
How would it be tested?
- Thirty to fifty publishable cases covering wrong OEC, drifting denominators, selection bias, guardrail conflicts, and insufficient information.
- Isolate development and test by scenario family, so a variant of a known case cannot leak across the split.
- Compare three arms: a fixed human checklist, a generic prompt, and the skeptic prompt.
- Score blind on recall of the decision-changing issue, rate of unfounded criticism, evidence faithfulness, actionability, and whether the correct decision changed.
- Bootstrap by case, and keep the author of the gold standard out of the scoring.
Next
What happens next?
Finish the five-user trial. If people reuse it and it beats a plain checklist, build the frozen benchmark and the comparison report. If they do not, it stays a prompt — which is a perfectly good outcome for a prompt.
Limitations
What this case does not establish.
- Nothing is deployed. Describing this as an agent platform would be false.
- A single author cannot write the gold standard and also score against it. Blind scoring needs someone else.
- "Users found it insightful" is a separate record from whether it was correct, and cannot substitute for it.