Fault-tolerant systems are designed to remain operational even when a part fails or experiences errors. This capability is crucial for maintaining services in critical applications and minimizing the impact of disruptions.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What you need to understand to reason about a domain.
Fault tolerance is a system's ability to stay functional and correct even when individual components fail or behave incorrectly. It usually relies on redundancy, clear fault boundaries, and mechanisms that keep service running through partial failures.
The concept grew out of practice in distributed and mission-critical IT, where hardware, networks, and software rarely fail completely and instead break in partial, uneven ways. Fault tolerance addresses that problem directly: a defect should not stop the whole service. Instead, systems isolate the fault, keep essential functions alive, and continue operation in a controlled way.
Think of a fault-tolerant system as a safety net with multiple supporting ropes. If one rope breaks, the others still carry the load. Monitoring notices deviations, redundant components take over, and recovery rules bring the service back to a defined state. The failure stays local instead of spreading through the whole system.
Multiple independent computers and services must work together reliably despite network and component failures.
Extra components or replicas provide alternatives when one part fails.
A defect stays confined to one part of the system so that problems do not spread.
Workload is switched to a standby component or another path when the active one fails.
After a failure, a defined state is rebuilt so the service can continue in a controlled way.
Fault tolerance matters for cloud services, platforms, payment systems, and other applications where downtime is expensive or risky. It improves availability and robustness, but it also adds infrastructure cost, operational overhead, and often more complexity around state and consistency. Perfect fault tolerance is rare; in practice, teams define a target level with explicit limits.
Where this building block is located in the topic model.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.