Alert Routing
Alert Routing and Escalation Policies¶
Alert routing and escalation policies in Prometheus Alertmanager define how alerts are directed to the right teams, channels, and individuals, and ensure critical issues are escalated promptly. These policies enable granular control over alert distribution, reducing noise and ensuring high-priority alerts are addressed before error budgets are exhausted.
Alert Routing Configuration¶
Alert routing determines which receivers (e.g., email, Slack, PagerDuty) receive alerts based on label selectors. Routes are defined in the route block of the Alertmanager configuration, with hierarchical matching rules.
Example: Team-Based Routing¶
route:
- name: "Production Team"
matchers:
- team = "production"
receivers:
- name: "slack-ops"
slack_configs:
- channel: "#alerts"
title: "Production Alert: {{ $labels.alertname }}"
Key Concepts¶
- Matchers: Use
match(exact) ormatch_re(regex) to filter alerts by labels (e.g.,team,severity). - Hierarchical Routing: Sub-routes inherit parent rules unless explicitly overridden.
- Default Receivers: Alerts not matching any route are sent to the default receiver.
Diagram: Alert Routing Flow¶
Escalation Policies¶
Escalation policies define how alerts are escalated to higher-priority recipients over time. This ensures alerts are not ignored and are resolved before they impact SLOs.
Example: Escalation Rule¶
escalation_config:
- name: "Escalate to Manager"
interval: 10m
receiver: "email-manager"
escalate_to:
- "slack-teamlead"
durations:
- 1h
Key Concepts¶
- Interval: How often the escalation check runs (e.g., every 10 minutes).
- Durations: Time thresholds for escalation (e.g., escalate after 1 hour if unresolved).
- Receiver Hierarchy: Escalations follow the same receiver hierarchy as routing.
Diagram: Escalation Over Time¶
[Alert Triggered]
↓
[Initial Receiver]
↓
[Escalation Check (10m)]
↓
[Escalate to Manager (after 1h)]
Best Practices¶
- Group Alerts by Team/Service: Use
group_byin routes to aggregate related alerts. - Test Routing Rules: Use the
--log.level=debugflag to simulate alert routing. - Monitor Escalation Paths: Track escalation metrics (e.g.,
alertmanager_escalation_total) to identify bottlenecks. - Avoid Over-Alerting: Use
forclauses in Prometheus rules to suppress transient alerts.
Example: Dynamic Receiver Configuration¶
route:
- name: "Dynamic Team Routing"
matchers:
- team = "dev"
receivers:
- name: "team-dev-slack"
slack_configs:
- channel: "#dev-team"
Key takeaways¶
- Use label matchers to route alerts to specific teams or services.
- Escalation policies ensure alerts are escalated to higher-priority recipients over time.
- Test routing and escalation rules with simulation tools like
alertmanager --config.file=... --log.level=debug. - Monitor routing and escalation metrics to optimize alert workflows.