Respond to incidents
Walk the incident-response flow end to end - detection, severity declaration, single-commander roles, containment through a runbook, communication, recovery, and the blameless postmortem - and see how business continuity and disaster recovery extend the same discipline to a region loss or a destructive event.
Units5
Duration29 min
Levelintermediate
By the end of this module, you'll be able to:
- Walk the incident lifecycle from detection through declaration, response, recovery, and the blameless postmortem, and name what enforces each step.
- Apply the single severity scale and the declare-high-downgrade-later principle, and explain why a security-relevant signal always overrides the apparent severity.
- Name the single-commander incident roles and the one rule that never bends: the Incident Commander never debugs hands-on and is never the same person as a responder.
- Explain how the BIA, service tiers, backup strategy, and DR plan chain together to meet an RTO/RPO, and in what order an AI-native org recovers when multiple services are down.
- Read Cogitave's ops/README.md as a worked example of an operational-resilience tree - the entry point to its incident-response and business-continuity docs - and follow it to the owning document for the detail.
Prerequisites
- Familiarity with how an AI-native org measures and reasons about reliability (SLIs, SLOs, error budgets) is recommended.
- No prior incident-response or on-call experience required.
Units
- 01Introduction4 min
- 02The incident-response model9 min
- 03Business continuity and disaster recovery9 min
- 04Knowledge check4 min
- 05Summary3 min
Related
- Operate from Day 1 Describe the Day 1 operating model - the pinned inner loop and the flow-based outer loop - state the parity contract that keeps the inner loop honest, read Cogitave's flow-based ways of working and its ADR-0025 choice of flow over Scrum as one org's answer, and explain why the human gate stays an exception inside the flow.
- Reliability and SRE Reason about a service's reliability targets the way an SRE discipline does - set an SLO from the user's journey, compute the error budget it buys, and operate the service against that budget with burn-rate alerts, sustainable on-call, a toil cap, and the freeze that stops shipping when the budget runs out.
- Run agentic operations Explain what agentic operations means - agents running an estate unattended inside rails - state the draft-vs-act rule that decides when an agent acts unattended versus when it stops for a human, and study Cogitave's own fleet of scheduled and operations agents as the worked example of the fleet you would build.