Operating processes describe recurring workflows, roles, and responsibilities for operating products and systems.
Operating processes provide stable, repeatable workflows for operating systems, services, and products. They define roles, responsibilities, escalation paths, and metrics for monitoring. They include procedures for deployments, monitoring, incident response, and change coordination aligned with business goals.
Time from detection to restoration of a service; measures response and recovery capability.
Share of deployments that cause failures or rollbacks; indicates release stability.
Percentage of time a service is available; relates to SLAs and SLOs.
SaaS companies use standardized runbooks for incident response and maintenance windows.
Mid-sized companies adopt ITIL elements for change and incident management to harmonize processes.
Teams adapt SRE principles for SLIs, SLOs, and error budgets to govern operating processes.
Take inventory of existing processes and tools.
Define roles, responsibilities, and escalation paths.
Create and validate runbooks for critical paths.
Automate repeatable steps and integrate into CI/CD.
Introduce metrics, dashboards, and regular reviews.