Menu

Earn Premium with Referrals

Invite your friends and earn Premium rewards through our referral program.

See how it works and start inviting friends.

Incident Response
HLD

Incident Response

When pages fire — roles, communication, and decisions under pressure until the fire is out.

The Lifecycle

 DETECT ──► TRIAGE ──► MITIGATE ──► RESOLVE ──► LEARN

 DETECT: alert/signal fires (user reports count too!)
 TRIAGE: severity assigned; roles claimed; channel opened
 MITIGATE: stop the bleeding FIRST (rollback > debugging)
 RESOLVE: root cause fixed; verification of recovery
 LEARN: blameless postmortem → actions tracked to DONE

 the discipline that separates mature orgs:
 MITIGATION BEFORE INVESTIGATION.
 a 10-minute rollback beats a 2-hour root-cause hunt while
 users suffer. diagnose AFTER stability returns.

Roles Under Pressure

 INCIDENT COMMANDER (IC):
   owns coordination + decisions, does NOT touch keyboards.
   declares incident, assigns work, calls rollbacks, closes it.

 OPERATIONS lead:
   hands-on investigation/mitigation. one at a time;
   spectators stay OUT of the critical path.

 COMMUNICATIONS lead:
   status updates to stakeholders on cadence:
     internal every 15-30min, external per policy.
   shields operators from interruption.

 small incidents: IC+ops may merge. big ones: NEVER let the
 person debugging also write stakeholder updates — both jobs
 fail quietly when combined.

Severity and Communication

SevDefinitionResponse
S1user-facing outage / data riskpage everyone, exec comms, all-hands
S2degraded, partial user impactimmediate on-call, business hours comms
S3minor/internalticket, next-business-day
 communication templates pre-written:
 "we are aware of X affecting Y; investigating; next update
  by HH:MM." — honest about unknowns, committed to cadence,
  no speculation about causes in public channels.

 timeline hygiene DURING the fight: a scribe logs actions +
   timestamps as they happen (postmortems live or die by this).

The Postmortem Contract

 BLAMELESS is operational, not decorative:
 humans act rationally given their context; fix CONTEXT not people.
 punishment guarantees the next failure hides longer.

 structure: impact (numbers!) · timeline · root cause(s) ·
 what went well/luckily · ACTION ITEMS each with owner+date.

 contributing-factor lens over single-root-cause:
   why did monitoring miss it? why did rollback take 40min?
   why did ONE bad config deploy to ALL regions?
 systemic answers prevent REPLAYS, not just reruns.

 close the loop: action items reviewed monthly until done —
 postmortems whose actions die in docs are theater.

Interview Framing

“Checkout is down, black Friday traffic incoming — walk me through your response” scored shape: lifecycle with mitigate-before-investigate emphasized, IC/ops/comms role split explained with WHY (cognitive load), severity table + comms-cadence template, blameless-postmortem mechanics with action-item tracking, scribe/timeline detail included. Incident questions test whether you’ve been IN the arena — role separation and rollback-first instincts are the tells.

My Private Notes

Notes are auto-saved locally to this device.