Recommendation Quality Beyond Engagement
When creator-side engagement improves, how do we confirm that viewers are actually getting better content?
Internal company work. This public case describes my contribution and the method; confidential definitions, data, and results are omitted.
- Problem
- A creator-side engagement goal was moving in the right direction. That number alone could not say whether viewers were getting better content, or the same reliable content more often.
- My contribution
- I defined consumption-side measures of the viewer experience, reported them by segment rather than in aggregate, and made the case for reading them next to the primary goal.
- What changed
- The creator-side metric stopped being read on its own. Viewer-experience measures and segment splits became part of the launch read.
The problem
What could the existing numbers not answer?
A short-form video feed has two sides that are easy to confuse. Creator-side engagement measures how much reaction the content collects. Viewer-side quality is whether the person scrolling got something worth their time. A ranking change can lift the first by leaning harder on what already works — showing the reliable creator again, resurfacing the format this viewer has reacted to before — without the second improving at all.
So the open question on a positive result was never "did the metric move." It was whether the metric had moved because the recommendations got better, or because they got more confident about a narrower set of content.
My contribution
What did I define, and why those measures?
I worked on the consumption side of the problem: measures that describe what the viewing session was actually like, reported in a form a launch review could act on.
- Repeat exposure
- How often a viewer meets near-duplicate or same-creator content inside a session and across sessions. A cheap way to lift engagement is to show the reliable thing again, and an aggregate engagement number is happy to let that happen.
- Viewing efficiency
- Time spent watching relative to the effort spent finding something worth watching. Engagement can rise while the search cost rises faster. The exact numerator and denominator are internal; the direction of the idea is what matters here.
- Segment reporting
- The same measures computed separately for viewer and creator segments rather than pooled. Averages across a heterogeneous population can move in a direction no individual segment experienced.
The integrity side of the work is a related but distinct measurement problem, which is why it is listed apart. Overall performance and performance on high-risk slices answer different questions. A sample sized for a population estimate does not automatically have the resolution to say anything about a rare, high-cost failure, so the evaluation sample has to be built for the decision it serves — with appropriate weighting when the population effect is what you need.
What a read looks like
How does a measure turn into an action?
Illustrative example. The structure below is how I organize this kind of read; the specific definitions, thresholds, and results from my work are not published here.
| Question the team has | What gets measured | Reported on | What the read triggers |
|---|---|---|---|
| Did engagement rise because supply narrowed? | Repeat exposure within and across sessions | Heavy and light viewers separately | A rise past the stated threshold holds the launch for review, even on a winning primary metric. |
| Is watching getting cheaper or more expensive? | Watch time relative to search effort | Same segments as the primary goal | A primary win with a falling efficiency reading is escalated as a tradeoff, not announced as a win. |
| Who is the average hiding? | The primary goal, computed per segment | Segments fixed before the read | Segment disagreement is reported in the launch note rather than filed as a follow-up analysis. |
A metric earns its place on a launch review by naming the action it triggers, not by being available.
The point of writing it this way is that each row ends in an action. A guardrail with no stated threshold and no consequence is a chart, and charts do not stop launches.
Method
What makes this kind of comparison credible?
These are the conditions I hold this work to. On a surface that changes continuously, an aggregate before-and-after cannot separate "ranking got better" from "ranking got more confident," so the design has to do that work.
- The experimental unit matches how the surface allocates content, so a viewer does not experience both arms across sessions.
- Consumption-side measures and their segments are named before the results are read. Heterogeneity found afterwards is a hypothesis, not a finding.
- Guardrails cover the failure the primary goal is structurally unable to see — here, repetition, narrowing supply, and integrity-relevant exposure.
- Each guardrail has a defined failure mode, a meaningful threshold, and enough sensitivity to detect harm at the size that would matter. A guardrail that never binds may mean the risk never materialized or the design already prevented it; what makes it decoration is having no threshold and no consequence attached.
Everything on this page is reported experience. The underlying data is internal and nothing here was re-run for this write-up.
What a technical reader can ask me
Where does the detail live?
- How a repeat-exposure measure gets defined in practice — the dedup key, the session boundary, and what breaks when either is wrong.
- How power is computed with repeated observations per viewer.
- How to size an evaluation sample so a rare, high-cost failure is actually resolvable, and how to weight back to a population estimate afterwards.
Limitations
What this case does not establish.
- Repeat exposure and viewing efficiency are proxies for experience. They are not satisfaction, and a viewer who is efficiently served mediocre content still had a mediocre session.
- This is my account of work done inside a team. Colleagues owned ranking, infrastructure, and shipping.
- Effect sizes, denominators, and time windows are not published here, so nothing on this page should be read as a quantified result.
- Reported experience
- Described from work I did inside a company. The underlying data is not public and is not reproduced here.