FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Saturday, 3 October 2026 · Oslo · London · New York

Enterprise · Long-form Report

Why Incident Postmortems Rarely Change Anything

Organisations write blameless postmortems, file them, and then relive the same outage. The document was never the mechanism. Without owners, deadlines and verification, a review is only an obituary.

Rows of dark server racks with a single corridor light, photographed in high-contrast black and white.
Rows of dark server racks with a single corridor light, photographed in high-contrast black and white.

Independent coverage

A

By Andrew Singer

Contributing Writer — AI / Data / Business · Freelance

Edited by Dr. Elin Lindqvist, MD

Published 1 October 2026

12 min read

Evidence: Analysis

Most engineering organisations now have a postmortem practice. Something breaks, an on-call engineer writes a timeline, the team meets, the discussion is blameless, a document is filed, and everyone returns to work. Six months later the same class of incident recurs, with different service names but the same shape: the same silent dependency, the same missing alert, the same manual step performed under pressure. The organisation did not lack analysis. It lacked a mechanism that turns analysis into change.

The gap is not cultural laziness. It is structural, and it exists because the postmortem is treated as a document when it is actually a project. A document has an author and an archive. A project has owners, deadlines, a budget, and a definition of done. The average postmortem has an author and an archive.

Action items without owners

Nearly every postmortem template ends with an action items section, and nearly every action item fails in the same place: at assignment. 'Improve alerting on the payments pipeline' is not an action item. It is a wish with punctuation. It has no named owner, no deadline, no definition of what would count as done, and no slot in any team's sprint plan. In organisations that track these things honestly, a large share of postmortem action items are never assigned to a person at all. They are assigned to a team, which is to say to no one.

The failure is compounded by scale. A severe incident generates fifteen or twenty action items, spread across five teams, none of whom attended the review. Each receiving team weighs the request against its existing commitments and quietly deprioritises it, because the incident is over and the roadmap was already full. Nobody says no. The item simply moves to the bottom of a backlog that is never emptied.

The review happens too far from the failure

There is also a timing problem. Reviews held one or two weeks after the incident are held in a different context: the on-call engineer has slept, the dashboards have been cleared, and the specific confusion of the outage has faded into a tidy narrative. Timelines reconstructed from memory flatten exactly the details that matter, the ambiguous signals that were misread, the runbook step that did not match reality, the escalation path that stalled in a channel nobody watches. The document records a clean story, and clean stories rarely contain the fix.

Aviation and other safety-critical fields learned this long ago. Investigations begin within hours, while physical evidence and human recollection are fresh, and the investigative bodies are permanent institutions with legal authority to demand information, not voluntary meetings scheduled around sprint planning. Software organisations copied the document format of incident investigation and skipped the institutional machinery that makes the format work.

The fixes that are never funded

Even well-assigned action items collide with a budgetary truth: the remediations that would actually prevent recurrence are usually structural, and structural work is never free. The incident happened because two services share a database, or because one manual step sits in the middle of an automated flow, or because the failover path has never been exercised. Fixing any of these takes weeks of engineering time that belongs to some product team's roadmap.

A common compromise emerges: the small, cheap action items get done, the structural ones do not. The team adds the missing alert and closes the ticket. But the alert is a detector, not a fix. It catches the next occurrence of the same failure and converts it from an outage into a page at three in the morning. The underlying fragility remains, and each future incident generates another round of detector work. Over the years the organisation accumulates a dense alarm system wrapped around an unchanged core.

Verification is the missing step

The deepest problem is that nobody verifies that action items worked. An alert added is not the same as an alert that fires correctly when the failure recurs. A runbook updated is not the same as a runbook that a sleep-deprived engineer can follow. An architecture change proposed is not an architecture shipped. The feedback loop closes only when the next incident of the same class occurs and the organisation checks, deliberately, whether the remediation held. Almost nobody performs this check, which means the organisation never learns whether its postmortem practice works at all.

The remedy is mundane. Every action item gets a named individual owner, a deadline, and an acceptance criterion. Items that require other teams' capacity get scheduled through the normal planning process, with the incident review as the sponsoring context. A recurring review, quarterly at most, audits open items from past incidents and marks which remediations have been validated against a real recurrence or a deliberate test. None of this is sophisticated. All of it is administration, which is precisely why it is skipped.

Blamelessness has its own failure mode

Blameless postmortems were a necessary correction to a culture of scapegoating, and they remain the right default. But a practice can drift from 'we do not punish individuals for systemic failures' to 'we do not examine decisions at all'. A review that treats every mistake as an inevitable consequence of circumstance produces documents that read like weather reports. Nobody could have done otherwise, therefore nothing should change, therefore the next incident will resemble this one.

The useful middle position distinguishes between the person and the decision. Individuals act under the constraints they face, and those constraints are the real subject of the review. But decisions made inside those constraints can still be examined: the decision to skip the load test, to defer the upgrade, to accept the single region. Examining decisions is not blame. It is the only route to changing the constraints that produced them.

What a working practice looks like

Organisations whose postmortems demonstrably reduce recurrence share a few visible habits. They convene the review within days, while the ambiguity is still recallable. They limit action items to a number that can actually be completed, and they would rather have five finished remediations than twenty open tickets. They track remediation work in the same system as feature work, with the same priority mechanics. They test their fixes, by injecting the failure again or exercising the failover, instead of assuming the ticket closing equals the risk closing. And they keep a visible register of recurring incident classes, so that a third occurrence of the same pattern triggers escalation rather than another routine review.

The point of a postmortem is not to understand the incident. Understanding is necessary and insufficient. The point is to make the next incident smaller, rarer, or faster to recover from, and that happens only through shipped, verified change. A postmortem that ends at the document has already failed. The document is the beginning of the work, not the record of its completion.

"A postmortem that ends at the document has already failed. The document is the beginning of the work, not the record of its completion."

Sources

Published 1 October 2026