A policy that defines a service's tolerable error budget and the organizational actions triggered when that budget is exceeded.
An Error Budget Policy specifies how much unreliability a service may tolerate over a defined period and which organizational actions trigger on breach. It ties SLOs to release, prioritization, and incident-response rules. The policy makes risk measurable and embeds accountability into operational governance.
Portion of time the SLO was met; central success control.
Speed at which the error budget is consumed; early warning indicator.
Average recovery time after an incident; measures responsiveness.
Google uses error budgets to balance reliability and feature velocity across services.
A startup defines simple SLOs and blocks releases when budgets are exceeded during peak periods.
Team prioritizes bug fixes over new features when the error burn rate reaches a critical level.
Define SLOs and SLIs for critical user journeys
Build monitoring pipelines and dashboards
Set policy rules for burn-rate thresholds and actions
Configure integrations to CI/CD and incident management