Timestamps are Required

For Both Executive and Technical Readers

A learning organization without a decision journal isn’t learning from experience, it’s learning from a reconstructed memory of experience, edited by the very outcome it should be testing.

Learning from a decision requires comparing what was predicted against what happened, and updating on the gap. That comparison needs two things: a record of the prediction, and a timestamp proving the record was made before the outcome was known. Without both, there is nothing to compare the outcome against, only a memory of what was predicted, and memory is reconstructive. Once an outcome is known, people unconsciously edit their recollection of what they expected, how confident they were, and why.

This is not a character flaw. It is how memory works, and it means the annual “lessons learned” review, run without dated records, is not an act of learning, it is an act of hindsight, dressed as one. An organization can hold that meeting every year for a decade and its judgment will not improve, because nothing in the process was ever falsifiable.

A decision journal is the fix, and it is deliberately narrow: four fields, recorded before the outcome, locked against later editing.

Decision Journal Entry recorded before the outcome is known
FieldWhat it captures
BeliefWhat you think is true right now, stated plainly enough to be judged right or wrong later
ConfidenceA number, not a feeling, so calibration can be scored, not just impression
DisconfirmerWhat specific evidence would change your mind, named in advance
Expected outcomeWhat you predict will happen, concrete enough to check
Timestamp: the field that makes the other four mean anything. An undated belief is just an opinion someone claims to have held.

The discipline is checking the record against reality later, not checking memory against reality. Most people, asked to recall a prediction they made six months ago, will unknowingly report something close to what happened, not what they said. The journal exists specifically to prevent that substitution.

Two well-documented biases do the damage, and they compound inside organizations in a way they do not for individuals.

BiasWhat it does
Hindsight bias“I knew it all along.” Once an outcome is known, people overestimate how predictable it was beforehand, and misremember their own prior uncertainty as having been lower than it was.
Outcome biasJudging the quality of a decision by how it turned out, rather than by the quality of the reasoning and the information available at the time. A good decision with a bad break gets marked wrong; a bad decision that got lucky gets marked right.

Inside an organization, these compound because the retelling happens socially, not privately. The executive who made the call gets to narrate what they “really thought” before anyone else’s memory can contradict them, and the retelling itself becomes the organization’s institutional memory. A team can be entirely well-intentioned and still end up with a shared history that is almost pure reconstruction, confident, detailed, and wrong.

Why This Is Worse for Teams Than Individuals An individual's hindsight bias distorts one person's memory. An organization's undocumented retrospective becomes the official record, repeated in the next planning cycle, cited as evidence the strategy worked, and passed to people who were not in the room to judge for themselves.

This is not a productivity habit; it is the mechanism behind every group that has been shown to improve its judgment over time.

Philip Tetlock’s Good Judgment Project identified “superforecasters” by this method: forecasters made dated, numerical predictions on real-world questions, the predictions were scored against what happened using a proper scoring rule, and the individuals who improved fastest were the ones whose calibration could be measured because their prior beliefs were locked in before the answer existed. There is no version of that finding that works without the timestamp: it is the whole mechanism, not a detail of the study design.

Annie Duke’s account of professional poker players makes the organizational version of the same point: the players who improve fastest separate the quality of a decision from the quality of its outcome, and the tool that makes that separation possible in practice is a written record made at the time of the decision, not a recollection of it afterward. “Resulting”, her term for outcome bias in this context, is the default failure mode without one.

The Causal-Inference Parallel This is structurally the same discipline Pearl’s twin-network method enforces for counterfactual reasoning: the factual branch has to be fixed, the exogenous terms recovered, before the counterfactual branch is evaluated, precisely so that knowledge of one branch cannot contaminate the abduction step for the other. A decision journal entry is the factual branch, written down before you know which way the twin network resolved. Reasoning about a past decision without one is exactly the error of conditioning on the outcome before doing the abduction, the causal-inference version of hindsight bias, not merely an analogy to it.

Put simply: calibration cannot be measured, and a decision cannot be evaluated on its merits, without an unforgeable record of what was believed before the fact. Every method that has been shown to work uses one. No method that skips it has been shown to work.

In practice, the question a stakeholder asks is rarely “show me the decision journal”, it’s some version of “did we make the right call,” and the answer depends entirely on whether a dated record exists to check against.

Looking back, did we make the right call raising prices in Q2?
AnyChatChecking the record, not memory
The Scope Card from that decision, dated April 14, recorded a 65% confidence that the increase would grow revenue, with the stated disconfirmer being “volume drop exceeds 3%.” Volume dropped 1.8%, within the range the model predicted. By the standard set before the decision, this was a good call, independent of how Q2 actually closed.
I remember being much more confident than that at the time.
AnyChatHindsight bias, named
That’s the pattern the dated record exists to catch. The 65% is what was signed off on before the outcome was known, the higher confidence is a reconstruction, formed after you learned the result. This isn’t a claim about which number is right, only about which one was made first.

This page names the principle behind artifacts that already exist elsewhere in the architecture, rather than introducing a new one. Every Scope Card and Elicitation Record is, structurally, a decision journal entry: a belief and a confidence level, dated and locked before the model is judged against reality. The Audit Record produced by every query is the same discipline applied automatically, at the pace of individual decisions rather than annual reviews. And the held-out log-score check in the Validation Report only means anything because the model’s predictions were generated, and recorded, before the held-out outcomes were revealed.

The uncomfortable version of this page’s argument: an organization that skips these artifacts is not choosing a lighter-weight process. It is choosing not to learn, while retaining the sincere belief that it does.