Observability and monitoring are crucial for understanding and managing complex systems.
A process for monitoring and notifying critical events.
A systematic approach to identifying and resolving incidents in IT environments.
Organized team duty to respond to incidents and operational disruptions outside regular hours. Purpose is rapid recovery, minimizing downtime, and providing clear escalation paths.
A policy that defines a service's tolerable error budget and the organizational actions triggered when that budget is exceeded.
A conceptual guide for systematically capturing, correlating and analysing telemetry (metrics, traces, logs) to enable fast debugging and performance optimisation.
A Service Level Objective (SLO) defines specific performance expectations for a service.
Strategic collection of telemetry from software and infrastructure to make behavior, performance and operational state measurable.
Concept for systematically collecting and forwarding metrics, logs and traces to support observability and operations.
Open standard and toolkit for instrumenting and collecting traces, metrics and logs via SDKs, collectors and exporters.
Technique for tracking and correlating requests across services to make performance issues and root causes in distributed systems visible.
Time-ordered records of events and state changes used for debugging, monitoring, and forensic analysis.
Metrics help measure and analyze the performance and efficiency of processes.
Systematic capture and visualization of dependencies between components, services and teams to support architecture and decision-making processes.
Technique for tracking and correlating requests across services to make performance issues and root causes in distributed systems visible.
Visual representation of services and their runtime dependencies to analyze communication, impact and failure sources.
Data visualization is the graphical representation of data to make patterns, trends, and insights visible.
Central dashboard for visualizing and analyzing telemetry (metrics, logs, traces) to enable rapid incident diagnosis and performance monitoring.
Grafana is an open-source tool for visualizing and analyzing data.