← All work

Working prototypeIndependent build

Narrative Intelligence: Evaluating Evolving Claims

How do you turn continuously changing opinions into a record that is traceable, comparable, and honest about uncertainty?

The pipeline runs over licensed sources in a private repository. The fixture linked below is written by me and is public, so the method can be checked without any access.

Problem
Opinions about a company arrive continuously, restate each other, and occasionally reverse. Reading them as a stream loses the one thing that matters: whether the underlying claim actually changed.
My contribution
I built the whole system — ingest, claim extraction, narrative clustering, stage and stance tracking, and the evaluation design that will say whether its judgements are any good.
Current result
The pipeline runs end to end and produces output. A public fixture is now checkable; measured quality is not, and the evaluation is designed rather than done.
My role
Sole author. Data pipeline, storage, narrative model, product surface, and evaluation design.
Methods
Claim extraction · Narrative clustering · Change detection · Timestamp discipline · Evaluation design
Evidence
Observational · Illustrative example
Also known as
Private repository over licensed sources
01

The task

What does the system actually do?

Given a stream of commentary about a company, produce a record of the distinct claims being made, group the ones that are actually the same claim, and detect when a claim changes stage or reverses. The output is meant to answer "what is new?" rather than "what was published?"

Built and running: a ticker-to-narrative pipeline with events and stages, stance-flip detection, content-hash caching so unchanged sources are not reprocessed, object-store ingest, and automated narrative ingest.

Not built: any systematic quality measurement. That is the honest state of this project and the reason it is labelled a working prototype rather than a finished case.

02

A public example you can check

What does one unit of input and expected output look like?

The linked fixture is six short passages I wrote by hand, with the grouping I believe is correct and the reason for each decision. It exists so a reader can disagree with my labels without needing the repository, a subscription, or my word for anything.

PassageClaimSame narrative?
AThe company benefits from data-centre capital spending by AI buyers.Grouped with the passage restating it in different words.
BThe company’s consumer graphics inventory has normalized after a glut.Kept separate. Same company, different mechanism, different time horizon.

Two of the six fixture passages. Shared vocabulary is the trap: both mention the same company and the same industry, and they are not the same claim.

The fixture contains inputs and my gold labels only. It does not contain model output or an error rate, because I have not run the annotation yet and will not publish a score I have not computed.

03

What the build has taught me

What goes wrong, and how confident am I about it?

Two failure classes show up repeatedly on inspection. Over-merging: two genuinely different theses about the same company collapse into one narrative because they share vocabulary. Retroactive coherence: once a stance flip is detected, the surrounding evidence reads as though it always pointed that way.

Neither is quantified. Naming a failure from inspection is not the same as measuring its rate, and I am not going to report an over-merge number I have not computed.

04

The evaluation this needs

How would the study be run so the result means something?

LayerEvaluation taskWhat must not be conflated
ExtractionAre entities, stances, and cited evidence faithful to the source?A fluent summary is not a correct extraction.
ClusteringAre same-narrative items merged and different ones kept apart?Shared vocabulary is not the same proposition.
TransitionIs a stage change or stance flip supported by new evidence?An explanation written after the price moved was not a signal at the time.
DigestDoes an update carry new information and reduce re-reading?Updating often is not decision value.
Market alignmentHow do timestamped claim changes relate to price changes?Price is an external outcome, not a label for extraction fidelity.

The evaluation study this project needs, and the confusion each layer prevents.

  • Start at roughly fifty closely read cases, then widen. This is proposed work, not an existing dataset.
  • Record published_at, available_at, ingested_at, and evaluated_at separately, freeze the inputs, and only then validate on a forward window. Without the four timestamps there is no way to show the model was not reading the future.
  • Hold out by narrative family and by source, so a near-duplicate article cannot appear on both sides of the split.
  • Compare deterministic rules, an LLM, and a hybrid on the same set, and report cost, latency, and error type rather than assuming the model wins.
  • Let annotators mark a stage as ambiguous. Forcing every passage into a definite stage manufactures agreement that does not exist.

Reposts are used for source diversity and disagreement, never as independent corroboration. A repost is not a second source, and counting it as one would inflate every confidence number in the system.

05

Next

What happens next, and what would stop it?

  • Extend the public fixture to cover hold, add, reverse, retract, and duplicated-source cases.
  • Run extraction and pairwise-clustering annotation on it, and publish the agreement rates including the bad ones.
  • Only then decide whether the product surface should expand or narrow.

Stop condition: if humans cannot agree on stage boundaries at a usable rate, the stage taxonomy is wrong and gets rewritten before any model is tuned against it.

06

Limitations

What this case does not establish.

  • The repository is private and sits on paid sources, so it is not a self-serve demo. The public fixture is the fix, and it is deliberately small.
  • The system produces output. That output has not been scored against human labels yet — the study is designed, not run.
  • Price movement is an external outcome. It is not a label for whether a claim was faithfully extracted, and I do not use it as one.
Observational
Measured on data that was not randomized. Supports hypotheses, not causal claims.
Illustrative example
Built on self-authored material to show the shape of a method. Not a result.