Writing3 Min Read
Writing Down What You Didn't Measure

Writing Down What You Didn't Measure

Aadesh Ingle
Try it yourselfClaim provenance ledger

I keep a small file in each project with a claim beside its evidence. "Extraction accuracy around 97%" points to the eval run. "The bad negotiation behavior cost real money" is marked as an estimate from affected threads because we never measured it cleanly. I started the file after noticing how my memory treated those two statements over time.

What memory does, given six months, is promote estimates into facts. A number you back-of-enveloped in a meeting gets repeated, and repetition launders it. By the third retelling, "roughly 40 of 50 in a quick harness we built that week" has become "80% reliability improvement," and you're not even lying. You've just lost the metadata. The claim outlived its uncertainty, and uncertainty is the part that made the claim honest.

The discipline

I record the measurement status when a claim first enters a deck, document, or client conversation, while I still remember how we obtained it.

It costs minutes. When a number enters a deck, a doc, or a client conversation, it gets a line in the file with one of three labels: measured (with a pointer to the run), estimated (with the method, however crude), or believed (the honest word for "experienced people think so"). That third category isn't shameful. Most operating knowledge lives there. The shame is only in dressing it as the first category.

Claims ledger

Follow a claim back to the evidence it actually has

Explore the article’s measured and estimated examples, and its category for experienced judgment. These evidence types do not form a sequential ladder.

Scenario

Existing article example: extraction accuracy around 97%; no new benchmark is asserted.

Sequence

  1. Keep the claimCurrent
  2. Label it measuredUpcoming
  3. Inspect its scopeUpcoming
Claim provenance: three separate branches for measured, estimated, and believed evidenceChoose the evidence label that is true todayWhat would make this claim auditable?Three evidence types · no promotion by repetition
Claim to repeatAccuracy around 97%
MeasuredRun + dataset + scoring
EstimatedMethod + uncertainty
BelievedExperience + context
Inspect the runScope, date, owner
Inspect the methodWhat cannot be recovered
Plan a measurementTest the workflow judgment

Keep the claim

“Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.

Read the full explanation

Extraction accuracy

Existing article example: extraction accuracy around 97%; no new benchmark is asserted.

  1. Keep the claim. “Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.
  2. Label it measured. Attach the eval run, dataset version, date, owner, and scoring method; the number is scoped to that measurement.
  3. Inspect its scope. A reader should be able to reconstruct what was scored. Do not generalize the article’s example to every production input.

Negotiation impact

Existing article example: a financial impact reconstructed from affected threads.

  1. Keep the claim. “The negotiation bug cost real money” is an estimated impact in the article, not a cleanly measured loss.
  2. Label it estimated. Record affected threads, the margin estimate method, and uncertainty alongside the claim.
  3. Preserve the limit. The article says clean measurement is no longer possible. Keep that limitation visible instead of promoting the estimate by repetition.

Experienced judgment

The article’s category for operating knowledge that has not been measured.

  1. Keep the claim. The article describes believed claims as judgments experienced people hold. Keep that status explicit instead of inventing a measured effect.
  2. Label it believed. Preserve the user conversations and workflow judgment that support the belief, with the absence of a measurement explicit.
  3. Name what to learn. Name the cheapest relevant observation or measurement that could test the operating judgment. New evidence may revise the claim; labels are not a sequential ladder.

Extraction accuracy

Existing article example: extraction accuracy around 97%; no new benchmark is asserted.

  1. Keep the claim. “Extraction accuracy around 97%” is the article’s example. The label is only useful alongside its evidence pointer.
  2. Label it measured. Attach the eval run, dataset version, date, owner, and scoring method; the number is scoped to that measurement.
  3. Inspect its scope. A reader should be able to reconstruct what was scored. Do not generalize the article’s example to every production input.

Negotiation impact

Existing article example: a financial impact reconstructed from affected threads.

  1. Keep the claim. “The negotiation bug cost real money” is an estimated impact in the article, not a cleanly measured loss.
  2. Label it estimated. Record affected threads, the margin estimate method, and uncertainty alongside the claim.
  3. Preserve the limit. The article says clean measurement is no longer possible. Keep that limitation visible instead of promoting the estimate by repetition.

Experienced judgment

The article’s category for operating knowledge that has not been measured.

  1. Keep the claim. The article describes believed claims as judgments experienced people hold. Keep that status explicit instead of inventing a measured effect.
  2. Label it believed. Preserve the user conversations and workflow judgment that support the belief, with the absence of a measurement explicit.
  3. Name what to learn. Name the cheapest relevant observation or measurement that could test the operating judgment. New evidence may revise the claim; labels are not a sequential ladder.

1 / 3 · Keep the claim

Speed
TakeawayKeep the exact claim, its evidence type, and its limits together so retelling cannot erase the uncertainty.
Try it yourself

Build a ledger row that keeps its uncertainty

Choose an evidence category and assemble a hypothetical record. Checking a field says what your draft would include; it does not create source evidence or verify the article’s claims.

Evidence category
Separate categorymeasuredRun and scoring
Selected categoryestimatedMethod and uncertainty
Separate categorybelievedExperience and context
Claim to preserveThe bad negotiation behavior cost real money

The article describes an estimate from affected threads and says clean measurement is no longer recoverable.

Include in the hypothetical evidence record
Draft metadata fields0 / 4Completeness is not a truth score.
Evidence category remainsestimatedRepetition and metadata do not promote it.
The draft still needs: Affected-thread evidence; Estimation method; Uncertainty and limits; Date and owner.
Inspect what the audit can establish

This exercise checks whether the category’s required metadata is represented. It cannot establish that a run exists, that an estimate is sound, or that an observation generalizes. Measured, estimated, and believed remain separate descriptions of the evidence available today.

Claim rowThe metadata that keeps a number from drifting

Claim: the exact sentence people are likely to repeat.

Status: measured, estimated, or believed. No fourth category called "basically true."

Evidence pointer: eval run, query, dashboard, transcript set, meeting note, or the absence of one.

Upgrade path: the cheapest measurement that would move the claim up one level.

The row is useful only if someone can audit it without asking you what you meant six months ago.

The file helps in three recurring situations:

  • Future-you can tell which numbers are load-bearing. When someone proposes building on top of "97%", you know instantly whether that foundation is concrete or hope.
  • Claims survive scrutiny because they arrive pre-scrutinized. When a client or interviewer pushes on a number and the answer is "that one's an estimate from affected threads, here's how we got it," the conversation gets more trusting, not less. People can hear the difference between calibration and confidence. They've all been burned by confidence.
  • It tells you what to measure next. The estimated column is a ranked backlog. The claims you repeat most often and know least firmly are, almost by definition, your most urgent measurement debt.

Against the demo voice

The pressure runs the other way, of course. The demo voice, the pitch voice, the interview voice all reward rounding up: tighter numbers, cleaner stories, the hedge sanded off. And in the short term it works, which is what makes it corrosive. Every inflated claim that lands successfully raises the baseline for the next one. You end up maintaining a reputation instead of having one.

"I don't know yet" and "we never measured that cleanly" can slow a conversation that wants certainty. Over time, marking estimates consistently makes a measured claim easier for other people to trust and reuse.

The practice requires only a small file and the discipline to update it while the evidence is still easy to find. That is enough to stop an estimate from acquiring false precision through repetition.

Source trailWhat this post is in conversation with
  • Service Level ObjectivesSRE service-level objectives show why a measurement needs a stable definition before a team can use it to make decisions.
  • Documenting Architecture DecisionsSame habit in architecture form: preserve context and consequences before memory rewrites why a decision happened.
  • DORA metricsDORA provides an example of operational claims tied to explicit metric definitions and tracked consistently over time.
End of entry

Keep track of what you have read.

Discussion