0votes
Incident response first 30 minutesWorkflowHigh risk
A calm, ordered playbook for when production is down.
We have an incident: {{symptom}} started at {{start_time}}. Guide me through the first 30 minutes: 1. Triage: questions to establish scope (who's affected, since when, what changed). 2. Stabilize: safest mitigations in order (roll back last deploy, disable feature flag, scale up, fail over). 3. Communicate: a status-page update and an internal update I can paste now. 4. Investigate: which dashboards, logs and recent changes to check, in order. 5. Record: what to write down for the postmortem while it's fresh. Ask me for the information you need step by step instead of guessing.

Log in to join the discussion.