c/devops › command
0votes

Incident response first 30 minutesWorkflowHigh risk

submitted by u/client to c/devops · 0 copies

A calm, ordered playbook for when production is down.

Workflow · 2 variables
We have an incident: {{symptom}} started at {{start_time}}.

Guide me through the first 30 minutes:
1. Triage: questions to establish scope (who's affected, since when, what changed).
2. Stabilize: safest mitigations in order (roll back last deploy, disable feature flag, scale up, fail over).
3. Communicate: a status-page update and an internal update I can paste now.
4. Investigate: which dashboards, logs and recent changes to check, in order.
5. Record: what to write down for the postmortem while it's fresh.

Ask me for the information you need step by step instead of guessing.
0 commentsreport
Sponsored
0 comments

Log in to join the discussion.