← All work

Case studyShort-form video platform

Recommendation Quality Beyond Engagement

When creator-side engagement improves, how do we confirm that viewers are actually getting better content?

Internal company work. This public case describes my contribution and the method; confidential definitions, data, and results are omitted.

Problem
A creator-side engagement goal was moving in the right direction. That number alone could not say whether viewers were getting better content, or the same reliable content more often.
My contribution
I defined consumption-side measures of the viewer experience, reported them by segment rather than in aggregate, and made the case for reading them next to the primary goal.
What changed
The creator-side metric stopped being read on its own. Viewer-experience measures and segment splits became part of the launch read.
My role
Measurement, analysis, and influencing the product decision. I did not build the ranking system.
Methods
Metric definition · Experiment design · Guardrail analysis · Segment heterogeneity
Evidence
Reported experience
01

The problem

What could the existing numbers not answer?

A short-form video feed has two sides that are easy to confuse. Creator-side engagement measures how much reaction the content collects. Viewer-side quality is whether the person scrolling got something worth their time. A ranking change can lift the first by leaning harder on what already works — showing the reliable creator again, resurfacing the format this viewer has reacted to before — without the second improving at all.

So the open question on a positive result was never "did the metric move." It was whether the metric had moved because the recommendations got better, or because they got more confident about a narrower set of content.

02

My contribution

What did I define, and why those measures?

I worked on the consumption side of the problem: measures that describe what the viewing session was actually like, reported in a form a launch review could act on.

Repeat exposure
How often a viewer meets near-duplicate or same-creator content inside a session and across sessions. A cheap way to lift engagement is to show the reliable thing again, and an aggregate engagement number is happy to let that happen.
Viewing efficiency
Time spent watching relative to the effort spent finding something worth watching. Engagement can rise while the search cost rises faster. The exact numerator and denominator are internal; the direction of the idea is what matters here.
Segment reporting
The same measures computed separately for viewer and creator segments rather than pooled. Averages across a heterogeneous population can move in a direction no individual segment experienced.

The integrity side of the work is a related but distinct measurement problem, which is why it is listed apart. Overall performance and performance on high-risk slices answer different questions. A sample sized for a population estimate does not automatically have the resolution to say anything about a rare, high-cost failure, so the evaluation sample has to be built for the decision it serves — with appropriate weighting when the population effect is what you need.

03

What a read looks like

How does a measure turn into an action?

Illustrative example. The structure below is how I organize this kind of read; the specific definitions, thresholds, and results from my work are not published here.

Question the team hasWhat gets measuredReported onWhat the read triggers
Did engagement rise because supply narrowed?Repeat exposure within and across sessionsHeavy and light viewers separatelyA rise past the stated threshold holds the launch for review, even on a winning primary metric.
Is watching getting cheaper or more expensive?Watch time relative to search effortSame segments as the primary goalA primary win with a falling efficiency reading is escalated as a tradeoff, not announced as a win.
Who is the average hiding?The primary goal, computed per segmentSegments fixed before the readSegment disagreement is reported in the launch note rather than filed as a follow-up analysis.

A metric earns its place on a launch review by naming the action it triggers, not by being available.

The point of writing it this way is that each row ends in an action. A guardrail with no stated threshold and no consequence is a chart, and charts do not stop launches.

04

Method

What makes this kind of comparison credible?

These are the conditions I hold this work to. On a surface that changes continuously, an aggregate before-and-after cannot separate "ranking got better" from "ranking got more confident," so the design has to do that work.

  • The experimental unit matches how the surface allocates content, so a viewer does not experience both arms across sessions.
  • Consumption-side measures and their segments are named before the results are read. Heterogeneity found afterwards is a hypothesis, not a finding.
  • Guardrails cover the failure the primary goal is structurally unable to see — here, repetition, narrowing supply, and integrity-relevant exposure.
  • Each guardrail has a defined failure mode, a meaningful threshold, and enough sensitivity to detect harm at the size that would matter. A guardrail that never binds may mean the risk never materialized or the design already prevented it; what makes it decoration is having no threshold and no consequence attached.

Everything on this page is reported experience. The underlying data is internal and nothing here was re-run for this write-up.

05

What a technical reader can ask me

Where does the detail live?

  • How a repeat-exposure measure gets defined in practice — the dedup key, the session boundary, and what breaks when either is wrong.
  • How power is computed with repeated observations per viewer.
  • How to size an evaluation sample so a rare, high-cost failure is actually resolvable, and how to weight back to a population estimate afterwards.
06

Limitations

What this case does not establish.

  • Repeat exposure and viewing efficiency are proxies for experience. They are not satisfaction, and a viewer who is efficiently served mediocre content still had a mediocre session.
  • This is my account of work done inside a team. Colleagues owned ranking, infrastructure, and shipping.
  • Effect sizes, denominators, and time windows are not published here, so nothing on this page should be read as a quantified result.
Reported experience
Described from work I did inside a company. The underlying data is not public and is not reproduced here.