Skip to content

SLIs & SLOs Overview

SLIs, SLOs, and Error Budgets Overview

Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets form the cornerstone of SRE practices, enabling teams to quantify reliability, set measurable goals, and balance innovation with stability. SLIs are objective metrics that reflect the performance of a service (e.g., latency, error rate, or traffic). SLOs are the targets for these metrics, defining the level of reliability users expect. Error budgets represent the allowable deviation from these targets, acting as a buffer to support innovation while maintaining service quality. Together, they provide a framework for managing reliability, prioritizing work, and communicating risks to stakeholders.


## Service Level Indicators (SLIs)

SLIs are the measurable, objective metrics used to assess the performance of a service. They must be specific, measurable, and actionable. Common examples include:

  • Latency: Time taken to respond to a request (e.g., 95th percentile response time).
  • Error Rate: Percentage of failed requests (e.g., 500/404 errors).
  • Traffic: Volume of requests served (e.g., requests per second).
  • Availability: Percentage of time the service is operational.

Example: A web service might use latency as an SLI to monitor how quickly it responds to user requests.

Command Example:

# Query latency metric in Prometheus  
curl http://localhost:9090/api/v1/query?query=latency_metric{job="my_service"}

Diagram:

[Service] --> [SLI Metrics]  
           |  
           v  
[Latency] [Error Rate] [Traffic]  


## Service Level Objectives (SLOs)

SLOs are the targets defined for SLIs, representing the level of reliability users expect. They are typically expressed as percentages and time-bound. For example:
- "99.9% availability over a 30-day period"
- "95% of requests must complete within 500ms"

SLOs are derived from SLIs and must be realistic, measurable, and aligned with business needs. They serve as the baseline for reliability and guide decisions about system improvements.

Example: A SLO might state, "Our API must have 99.9% availability for all users, with no more than 0.1% downtime per month."

Command Example:

# Calculate SLO compliance using Prometheus  
curl http://localhost:9090/api/v1/query?query=avg_over_time(count_requests{status="200"}[7d]) / avg_over_time(count_requests_total[7d])


## Error Budgets

Error budgets are the allowable deviation from SLO targets, calculated as:

Error Budget = (1 - SLO Percentage) × Total Time Period  
For example, an SLO of 99.9% availability over a 30-day period allows 0.1% (approximately 43.2 minutes) of downtime.

Error budgets enable teams to balance reliability and innovation. When the budget is not exhausted, teams can prioritize new features or optimizations. When the budget is breached, it signals the need to investigate and resolve issues.

Example: If a service exceeds its error budget, the team might deploy a rollback or investigate root causes.

Command Example:

# Calculate error budget usage in Prometheus  
curl http://localhost:9090/api/v1/query?query=1 - (avg_over_time(count_requests{status="200"}[7d]) / avg_over_time(count_requests_total[7d]))


Key takeaways

  • SLIs are the metrics that quantify service performance (e.g., latency, error rate).
  • SLOs are the targets for SLIs, defining acceptable levels of reliability.
  • Error budgets act as a buffer, allowing teams to innovate while maintaining service quality.
  • Regularly monitor SLIs, track SLO compliance, and use error budgets to prioritize reliability efforts.
  • Tools like Prometheus and Grafana are essential for visualizing and analyzing SLI/SLO data.