Strategies, processes and technical measures to restore IT systems and data after major outages or disasters.
Disaster recovery defines strategies, processes and technologies to restore IT systems and data after major outages. The goal is to minimize downtime (RTO) and data loss (RPO) through planning, backups, replication and tested recovery procedures. It covers organizational processes, technical measures and regular validation tests.
Maximum tolerable time to restore a service.
Maximum tolerable data loss in time (time since last consistent backup).
Average time required to fully recover after an outage.
Annual DR exercise with failover to a secondary site and measurement of service recovery times.
Recovery after an encryption attack using point-in-time backups and rebuild of affected systems.
Automated failover between cloud regions to minimize visible downtime for customers.
Analyze critical services and define RTO/RPO.
Design and implement redundancy and replication architecture.
Create runbooks and automation scripts.
Perform regular tests and drills and adapt processes.