CogitaveLearn

Reliability and SRE

Reason about a service's reliability targets the way the reliability standard does - set an SLO from the user's journey, compute the error budget it buys, and operate the service against that budget with burn-rate alerts, sustainable on-call, a toil cap, and the freeze that stops shipping when the budget runs out.

Units5
Duration29 min
Levelintermediate

By the end of this module, you'll be able to:

  • Explain why reliability is engineered as a budgeted feature, not chased as an absolute, and why SRE is an engineering discipline rather than a renamed ops rota.
  • Set an SLO from a critical user journey, pick the right tier target, and compute the error budget and burn rate it produces.
  • State the error-budget policy - freeze on exhaustion, one named owner, security exempt - and say what it buys operationally: sustainable on-call, a capped toil load, tested resilience, and an SLA that is always looser than the SLO behind it.
  • Read the reliability standard and ADR-0020 as the source of truth for a service's reliability posture, and know when a different value requires an Accepted RFC.

Prerequisites

  • Completion of "Work the request lifecycle", or equivalent familiarity with the estate's propose-only change process.
  • Familiarity with SLIs and telemetry is helpful; this module defines SLO and error budget from first principles.

Units

  1. 01Introduction4 min
  2. 02SLOs and error budgets8 min
  3. 03Operating for reliability9 min
  4. 04Knowledge check5 min
  5. 05Summary3 min