Skip to content

Postmortem Workflow Runbook

Use this runbook after service recovery is verified and before final incident closure. It turns the incident timeline, decisions, actions, and recovery evidence into an approved record with measurable follow-up.

A postmortem explains what the evidence supports. Keep unsupported causes as hypotheses and preserve machine records rather than rewriting history to make the timeline cleaner.

Start when service recovery has passed its observation window and the incident owner can identify the authoritative audit and evidence references.

RoleResponsibility
Incident ownerOwns impact, chronology, recovery status, and closure decision
FacilitatorLeads review and separates evidence from interpretation
Action ownerAccepts a follow-up with a due date and measurable completion evidence
ReviewerConfirms the record is supported, complete, and safe to share

The facilitator should not use the review to assign personal blame. The unit of analysis is the system, decision context, and control that did or did not work.

  • Incident record: scope, severity history, members, owners, and state transitions.
  • Audit trail: detected issues, decisions, approvals, actions, no-ops, retries, and rollback.
  • Evidence set: metrics, logs, traces, changes, notifications, and cited knowledge.
  • Impact record: affected capability, duration, population, and SLO effect.
  • Recovery proof: restored state, verification window, and residual risk.

Do not start from chat recollection when authoritative records exist. Recollection can add context but should be labeled as a participant statement.

  1. Create the draft. Generate an initial chronology from incident and append-only audit records without changing their timestamps or content.
  2. Verify impact. Confirm the start, detection, mitigation, recovery, and end times, plus the affected capability and measured SLO effect.
  3. Reconstruct decisions. For each key decision, record the evidence available at that time, the selected branch, and the resulting outcome.
  4. Separate causes. Distinguish root cause, contributing conditions, detection gaps, response gaps, and recovery gaps.
  5. Test claims. Link each causal statement to evidence and retain credible alternatives that were not disproved.
  6. Assess controls. Record which detector, rule, approval, stop condition, rollback, notification, and audit controls worked or failed.
  7. Define follow-up. Create corrective and preventive actions with owners, due dates, priority, and measurable completion evidence.
  8. Review and approve. Resolve unsupported claims, confirm sensitive data is excluded, and obtain the required reviewer approval.
  9. Link and close. Attach the approved postmortem to the incident and close only after unresolved risk and action ownership are explicit.
CheckpointWhat to capture
First impactEarliest supported user or operation impact
DetectionFirst detected issue and when it reached a durable route
TriageSeverity, owner, scope, and first decision deadline
MitigationProposal, decision, approval, execution, and observed effect
Rollback or recoveryTrigger, action, verification, and residual impact
Stable serviceStart and end of the recovery observation window

A useful follow-up changes a control or closes an evidence gap. Avoid actions such as “be more careful” that have no owner or test.

Required fieldExample evidence target
Owner and due dateNamed accountable role and review date
Control changedRule, runbook, test, alert, rollback, or provider reference
Completion proofPassing scenario, drill record, or measured production signal
Safety modeObservation mode evidence before any enforcement change
Closure conditionObjective result that lets the action be closed

Reusable rule, runbook, or knowledge improvements remain inert candidates until their normal review and promotion path accepts them.

Do not close when impact, recovery, unresolved risk, owner, or required follow-up is missing. Pause review when the evidence set changes materially, timestamps conflict, a cited record cannot be verified, or sensitive data has entered the draft. Keep unsupported causes labeled as hypotheses.

The approved record should link the incident, evidence-set version, timeline, RCA claims, response actions, approvals, rollback, recovery proof, reviewers, and follow-up items. Record the postmortem version and approval in the incident audit trail.

Complete the workflow when the reviewed postmortem is linked, residual risk is accepted by the right owner, and every required follow-up has an owner, due date, and evidence target. Follow-up completion can occur after incident closure.

To continue withRead
Recheck source evidence and causal claimsRCA evidence collection
Improve a detector with frozen scenariosAlert tuning
Validate a corrective resilience controlChaos game day