DevOps and Platform Engineering are approaches aimed at improving collaboration between software development and IT operations.
This segment covers concepts and approaches for ensuring stable and observable system behavior during operation. It includes mechanisms for measuring and evaluating reliability, capturing states and events, and analyzing deviations and failures. It describes how platforms and services are made observable in order to understand and interpret their behavior over time, without focusing on infrastructure definition or delivery processes. Topics related to platform provisioning, software delivery, or business logic are addressed in other segments.
A structured approach to identify the root causes of problems.
Metrics help measure and analyze the performance and efficiency of processes.
Observability enables understanding the state of complex systems through metrics, logs, and traces.
Reliability is a critical concept in system development that ensures systems consistently perform as expected.
A Service Level Agreement (SLA) defines the expectations for the services provided by a vendor.
A Service Level Indicator (SLI) measures the quality of a service against predefined criteria.
A Service Level Objective (SLO) defines specific performance expectations for a service.
Grafana is an open-source tool for visualizing and analyzing data.
The ELK Stack combines Elasticsearch, Logstash, and Kibana for efficient data processing and visualization.
Prometheus is an open-source monitoring and alerting system.