Runbook 04 · Learning

Write the system change, not the courtroom transcript

A factual postmortem structure that preserves impact, decisions, and verifiable follow-up without rewarding blame or theater.

RB–04 / v1.0Output: Postmortem + action registerPrint-friendly

Use this for: material incidents, risky near-misses, repeated alerts, and recoveries where the team learned something the operating system should retain.

Ground rules

  • Describe what happened with evidence available at the time.
  • Do not use hindsight as proof that a decision was negligent.
  • Treat human action as part of the system: permissions, interfaces, incentives, fatigue, documentation, and staffing all shape it.
  • Publish only what is appropriate for the audience; remove secrets and personal or customer data.

The document

1. Executive summary

Two or three sentences: the user impact, duration, service involved, and how normal operation was restored.

2. Impact

State affected users, failed journeys, data consequences, contractual or security implications, and the evidence behind estimates. Avoid invented precision.

3. Timeline

Use UTC. Include detection, decisions, actions, vendor contact, recovery, and verification. Separate observation from interpretation.

4. Detection and response

How did the team learn about the failure? What signal should have fired sooner? Which response steps helped or delayed recovery?

5. Contributing conditions

Explain the conditions that made the incident possible or harder to recover: architecture, defaults, ownership, access, capacity, release process, vendor behavior, or missing tests.

6. What changed

Record immediate mitigation separately from durable prevention. Not every incident needs a complex rebuild; every action needs an owner, due date, and verification method.

Copy-ready postmortem

Incident:
Date / duration (UTC):
Authors:
Status: DRAFT / REVIEWED / COMPLETE

SUMMARY
[User impact, duration, recovery]

IMPACT
[Scope and evidence]

TIMELINE (UTC)
TIME | OBSERVATION / DECISION / ACTION | EVIDENCE
-----|---------------------------------|---------
     |                                 |

DETECTION AND RESPONSE
[What signaled; what delayed or helped]

CONTRIBUTING CONDITIONS
[Technical and organizational conditions]

RECOVERY
[Why the chosen action worked; how verified]

ACTION | OWNER | DUE | VERIFICATION | STATUS
-------|-------|-----|--------------|-------
       |       |     |              |

WHAT WENT WELL
-

OPEN QUESTIONS
- 

Action quality test

  • Names one accountable owner, not “the team.”
  • Has a due date tied to risk.
  • Changes a system, guardrail, test, interface, or operating rule.
  • Defines observable evidence of completion.
  • Is closed only after verification, not after a ticket is created.

Same incident, different week?
Hoyack can review the service and its operating path.

Request a rescue review