Organized team duty to respond to incidents and operational disruptions outside regular hours. Purpose is rapid recovery, minimizing downtime, and providing clear escalation paths.
On-call describes the structured rotation of responsibilities for operations teams to quickly detect, escalate, and remediate incidents. It includes scheduling, alerting, runbooks and post-incident reviews for continuous improvement. Well-designed on-call processes reduce downtime and distribute operational knowledge across the organization.
Average time until an alert is acknowledged by on-call.
Average time until normal operation is restored after an incident.
Share of irrelevant or false-positive alerts relative to total alerts.
An SRE team runs rotating on-call, combining alert prioritization and runbooks for remediation.
A small product team shares on-call duties among few people and uses clear escalation paths.
A team uses a pager/incident service for alerting, complemented by automated playbooks.
Define goals, SLAs and compensation rules for on-call.
Introduce a rotating roster and clear escalation paths.
Create and test runbooks; introduce and monitor metrics.