The Lifecycle
DETECT ──► TRIAGE ──► MITIGATE ──► RESOLVE ──► LEARN
DETECT: alert/signal fires (user reports count too!)
TRIAGE: severity assigned; roles claimed; channel opened
MITIGATE: stop the bleeding FIRST (rollback > debugging)
RESOLVE: root cause fixed; verification of recovery
LEARN: blameless postmortem → actions tracked to DONE
the discipline that separates mature orgs:
MITIGATION BEFORE INVESTIGATION.
a 10-minute rollback beats a 2-hour root-cause hunt while
users suffer. diagnose AFTER stability returns.
Roles Under Pressure
INCIDENT COMMANDER (IC):
owns coordination + decisions, does NOT touch keyboards.
declares incident, assigns work, calls rollbacks, closes it.
OPERATIONS lead:
hands-on investigation/mitigation. one at a time;
spectators stay OUT of the critical path.
COMMUNICATIONS lead:
status updates to stakeholders on cadence:
internal every 15-30min, external per policy.
shields operators from interruption.
small incidents: IC+ops may merge. big ones: NEVER let the
person debugging also write stakeholder updates — both jobs
fail quietly when combined.
Severity and Communication
| Sev | Definition | Response |
|---|---|---|
| S1 | user-facing outage / data risk | page everyone, exec comms, all-hands |
| S2 | degraded, partial user impact | immediate on-call, business hours comms |
| S3 | minor/internal | ticket, next-business-day |
communication templates pre-written:
"we are aware of X affecting Y; investigating; next update
by HH:MM." — honest about unknowns, committed to cadence,
no speculation about causes in public channels.
timeline hygiene DURING the fight: a scribe logs actions +
timestamps as they happen (postmortems live or die by this).
The Postmortem Contract
BLAMELESS is operational, not decorative:
humans act rationally given their context; fix CONTEXT not people.
punishment guarantees the next failure hides longer.
structure: impact (numbers!) · timeline · root cause(s) ·
what went well/luckily · ACTION ITEMS each with owner+date.
contributing-factor lens over single-root-cause:
why did monitoring miss it? why did rollback take 40min?
why did ONE bad config deploy to ALL regions?
systemic answers prevent REPLAYS, not just reruns.
close the loop: action items reviewed monthly until done —
postmortems whose actions die in docs are theater.
Interview Framing
“Checkout is down, black Friday traffic incoming — walk me through your response” scored shape: lifecycle with mitigate-before-investigate emphasized, IC/ops/comms role split explained with WHY (cognitive load), severity table + comms-cadence template, blameless-postmortem mechanics with action-item tracking, scribe/timeline detail included. Incident questions test whether you’ve been IN the arena — role separation and rollback-first instincts are the tells.
Premium Content
Unlock Incident Response and all premium lessons with a subscription.
All premium lessons
Ad-free experience
Priority support
From ₹199.99/year — See plans