DevOps & Platform Engineering
DevOps and Platform Engineering are approaches aimed at improving collaboration between software development and IT operations.
- Knowledge domains
- /Thematic areas
- /Segments
- /Building blocks
Reliability & Observability
This segment covers concepts and approaches for ensuring stable and observable system behavior during operation. It includes mechanisms for measuring and evaluating reliability, capturing states and events, and analyzing deviations and failures. It describes how platforms and services are made observable in order to understand and interpret their behavior over time, without focusing on infrastructure definition or delivery processes. Topics related to platform provisioning, software delivery, or business logic are addressed in other segments.
Root Cause Analysis (RCA)
A structured approach to identify the root causes of problems.
Metrics
Metrics help measure and analyze the performance and efficiency of processes.
Observability
Observability enables understanding the state of complex systems through metrics, logs, and traces.
Reliability
Reliability is a critical concept in system development that ensures systems consistently perform as expected.
Service Level Agreement (SLA)
A Service Level Agreement (SLA) defines the expectations for the services provided by a vendor.
Service Level Indicator (SLI)
A Service Level Indicator (SLI) measures the quality of a service against predefined criteria.
Service Level Objective (SLO)
A Service Level Objective (SLO) defines specific performance expectations for a service.
Grafana
Grafana is an open-source tool for visualizing and analyzing data.
ELK Stack (Elasticsearch, Logstash, Kibana)
The ELK Stack combines Elasticsearch, Logstash, and Kibana for efficient data processing and visualization.
Prometheus
Prometheus is an open-source monitoring and alerting system.