Debrief
What has to happen between a first voice memo and a user who saves the result and comes back?
Live product with anonymous trial and early users beyond me. Retention and extraction accuracy are not measured yet, so no usage figures appear here.
- Problem
- A voice app can produce a summary quickly and still not produce anything a user trusts enough to keep. The interesting step is the one between those two.
- My contribution
- I built the product end to end — capture, transcription, classification, daily and weekly reviews, open-loop tracking — and defined what to measure first.
- Current result
- Live, with early users beyond me. The next step is understanding where first-time users find value, where extraction needs correction, and what brings them back.
The task
What does the product do, and what is measurable?
Debrief turns spoken notes into structured reflections and follow-up actions: anonymous trial, voice capture, transcription and classification, daily and weekly reviews, and open-loop tracking. A user can reach a result without an account, which makes the first-value moment easy to observe and the long-run behaviour harder to see.
- Time to value
- From first recording to a summary the user reads. If this is slow, nothing downstream matters.
- Extraction error rate
- How often a task or theme is invented, missed, or mis-assigned. Requires a labelled set, which does not exist yet.
- Anonymous-to-saved conversion
- The step where a user decides the output is worth keeping. This is the honest first metric for the product.
Where it stands
What is actually known?
The product is live and has early users beyond me. What I cannot yet say is where first-time users find value, how often extraction needs correcting, or what brings anyone back — none of that is instrumented well enough to report, and I would rather say so than publish a number I would not defend.
My own daily use is a design signal: it keeps the product honest about friction. It is not retention evidence and is not reported as such.
How this should be studied
What would make the next decision an informed one?
- Instrument the funnel to the save action before adding capability.
- Hand-label a small set of recordings for extracted tasks, then measure precision and recall against it.
- Separate "the transcription was wrong" from "the extraction was wrong" — they have different fixes.
- Keep the display name and the repository name explained in one place, so the product, the site, and the README agree.
Limitations
What this case does not establish.
- There are early external users, but no reliable retention or outcome measurement yet. I do not report user counts, activity, or anything resembling product-market fit.
- Task extraction errors are visible on inspection but have not been counted against a labelled set.
- The anonymous trial and the signed-in stage can each be measured; connecting them depends on instrumentation and identity rules that are not fully in place, which is a gap in what I currently measure rather than something unmeasurable.
- Observational
- Measured on data that was not randomized. Supports hypotheses, not causal claims.