I keep a small file in each project with a claim beside its evidence. "Extraction accuracy around 97%" points to the eval run. "The bad negotiation behavior cost real money" is marked as an estimate from affected threads because we never measured it cleanly. I started the file after noticing how my memory treated those two statements over time.
What memory does, given six months, is promote estimates into facts. A number you back-of-enveloped in a meeting gets repeated, and repetition launders it. By the third retelling, "roughly 40 of 50 in a quick harness we built that week" has become "80% reliability improvement," and you're not even lying. You've just lost the metadata. The claim outlived its uncertainty, and uncertainty is the part that made the claim honest.
The discipline
I record the measurement status when a claim first enters a deck, document, or client conversation, while I still remember how we obtained it.
It costs minutes. When a number enters a deck, a doc, or a client conversation, it gets a line in the file with one of three labels: measured (with a pointer to the run), estimated (with the method, however crude), or believed (the honest word for "experienced people think so"). That third category isn't shameful. Most operating knowledge lives there. The shame is only in dressing it as the first category.
Claims ledger
Follow a claim back to the evidence it actually has
Explore the article’s measured and estimated examples, and its category for experienced judgment. These evidence types do not form a sequential ladder.
Scenario
Existing article example: extraction accuracy around 97%; no new benchmark is asserted.
Sequence
- Keep the claimCurrent
- Label it measuredUpcoming
- Inspect its scopeUpcoming
Keep the claim
“Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.
Read the full explanation
Extraction accuracy
Existing article example: extraction accuracy around 97%; no new benchmark is asserted.
- Keep the claim. “Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.
- Label it measured. Attach the eval run, dataset version, date, owner, and scoring method; the number is scoped to that measurement.
- Inspect its scope. A reader should be able to reconstruct what was scored. Do not generalize the article’s example to every production input.
Negotiation impact
Existing article example: a financial impact reconstructed from affected threads.
- Keep the claim. “The negotiation bug cost real money” is an estimated impact in the article, not a cleanly measured loss.
- Label it estimated. Record affected threads, the margin estimate method, and uncertainty alongside the claim.
- Preserve the limit. The article says clean measurement is no longer possible. Keep that limitation visible instead of promoting the estimate by repetition.
Experienced judgment
The article’s category for operating knowledge that has not been measured.
- Keep the claim. The article describes believed claims as judgments experienced people hold. Keep that status explicit instead of inventing a measured effect.
- Label it believed. Preserve the user conversations and workflow judgment that support the belief, with the absence of a measurement explicit.
- Name what to learn. Name the cheapest relevant observation or measurement that could test the operating judgment. New evidence may revise the claim; labels are not a sequential ladder.
Extraction accuracy
Existing article example: extraction accuracy around 97%; no new benchmark is asserted.
- Keep the claim. “Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.
- Label it measured. Attach the eval run, dataset version, date, owner, and scoring method; the number is scoped to that measurement.
- Inspect its scope. A reader should be able to reconstruct what was scored. Do not generalize the article’s example to every production input.
Negotiation impact
Existing article example: a financial impact reconstructed from affected threads.
- Keep the claim. “The negotiation bug cost real money” is an estimated impact in the article, not a cleanly measured loss.
- Label it estimated. Record affected threads, the margin estimate method, and uncertainty alongside the claim.
- Preserve the limit. The article says clean measurement is no longer possible. Keep that limitation visible instead of promoting the estimate by repetition.
Experienced judgment
The article’s category for operating knowledge that has not been measured.
- Keep the claim. The article describes believed claims as judgments experienced people hold. Keep that status explicit instead of inventing a measured effect.
- Label it believed. Preserve the user conversations and workflow judgment that support the belief, with the absence of a measurement explicit.
- Name what to learn. Name the cheapest relevant observation or measurement that could test the operating judgment. New evidence may revise the claim; labels are not a sequential ladder.
Build a ledger row that keeps its uncertainty
Choose an evidence category and assemble a hypothetical record. Checking a field says what your draft would include; it does not create source evidence or verify the article’s claims.
The article describes an estimate from affected threads and says clean measurement is no longer recoverable.
Inspect what the audit can establish
This exercise checks whether the category’s required metadata is represented. It cannot establish that a run exists, that an estimate is sound, or that an observation generalizes. Measured, estimated, and believed remain separate descriptions of the evidence available today.
Claim: the exact sentence people are likely to repeat.
Status: measured, estimated, or believed. No fourth category called "basically true."
Evidence pointer: eval run, query, dashboard, transcript set, meeting note, or the absence of one.
Upgrade path: the cheapest measurement that would move the claim up one level.
The file helps in three recurring situations:
- Future-you can tell which numbers are load-bearing. When someone proposes building on top of "97%", you know instantly whether that foundation is concrete or hope.
- Claims survive scrutiny because they arrive pre-scrutinized. When a client or interviewer pushes on a number and the answer is "that one's an estimate from affected threads, here's how we got it," the conversation gets more trusting, not less. People can hear the difference between calibration and confidence. They've all been burned by confidence.
- It tells you what to measure next. The estimated column is a ranked backlog. The claims you repeat most often and know least firmly are, almost by definition, your most urgent measurement debt.
Against the demo voice
The pressure runs the other way, of course. The demo voice, the pitch voice, the interview voice all reward rounding up: tighter numbers, cleaner stories, the hedge sanded off. And in the short term it works, which is what makes it corrosive. Every inflated claim that lands successfully raises the baseline for the next one. You end up maintaining a reputation instead of having one.
"I don't know yet" and "we never measured that cleanly" can slow a conversation that wants certainty. Over time, marking estimates consistently makes a measured claim easier for other people to trust and reuse.
The practice requires only a small file and the discipline to update it while the evidence is still easy to find. That is enough to stop an estimate from acquiring false precision through repetition.
- Service Level ObjectivesSRE service-level objectives show why a measurement needs a stable definition before a team can use it to make decisions.
- Documenting Architecture DecisionsSame habit in architecture form: preserve context and consequences before memory rewrites why a decision happened.
- DORA metricsDORA provides an example of operational claims tied to explicit metric definitions and tracked consistently over time.

