Skip to content

Alert Routing

Alert Routing and Escalation Policies

Alert routing and escalation policies in Prometheus Alertmanager define how alerts are directed to the right teams, channels, and individuals, and ensure critical issues are escalated promptly. These policies enable granular control over alert distribution, reducing noise and ensuring high-priority alerts are addressed before error budgets are exhausted.


Alert Routing Configuration

Alert routing determines which receivers (e.g., email, Slack, PagerDuty) receive alerts based on label selectors. Routes are defined in the route block of the Alertmanager configuration, with hierarchical matching rules.

Example: Team-Based Routing

route:
  - name: "Production Team"
    matchers:
      - team = "production"
    receivers:
      - name: "slack-ops"
        slack_configs:
          - channel: "#alerts"
            title: "Production Alert: {{ $labels.alertname }}"

Key Concepts

  • Matchers: Use match (exact) or match_re (regex) to filter alerts by labels (e.g., team, severity).
  • Hierarchical Routing: Sub-routes inherit parent rules unless explicitly overridden.
  • Default Receivers: Alerts not matching any route are sent to the default receiver.

Diagram: Alert Routing Flow

[Alert Generated] 
  ↓
[Route Matchers] 
  ↓
[Receiver (e.g., Slack, Email)] 

Escalation Policies

Escalation policies define how alerts are escalated to higher-priority recipients over time. This ensures alerts are not ignored and are resolved before they impact SLOs.

Example: Escalation Rule

escalation_config:
  - name: "Escalate to Manager"
    interval: 10m
    receiver: "email-manager"
    escalate_to: 
      - "slack-teamlead"
    durations:
      - 1h

Key Concepts

  • Interval: How often the escalation check runs (e.g., every 10 minutes).
  • Durations: Time thresholds for escalation (e.g., escalate after 1 hour if unresolved).
  • Receiver Hierarchy: Escalations follow the same receiver hierarchy as routing.

Diagram: Escalation Over Time

[Alert Triggered] 
  ↓
[Initial Receiver] 
  ↓
[Escalation Check (10m)] 
  ↓
[Escalate to Manager (after 1h)] 

Best Practices

  1. Group Alerts by Team/Service: Use group_by in routes to aggregate related alerts.
  2. Test Routing Rules: Use the --log.level=debug flag to simulate alert routing.
  3. Monitor Escalation Paths: Track escalation metrics (e.g., alertmanager_escalation_total) to identify bottlenecks.
  4. Avoid Over-Alerting: Use for clauses in Prometheus rules to suppress transient alerts.

Example: Dynamic Receiver Configuration

route:
  - name: "Dynamic Team Routing"
    matchers:
      - team = "dev"
    receivers:
      - name: "team-dev-slack"
        slack_configs:
          - channel: "#dev-team"

Key takeaways

  • Use label matchers to route alerts to specific teams or services.
  • Escalation policies ensure alerts are escalated to higher-priority recipients over time.
  • Test routing and escalation rules with simulation tools like alertmanager --config.file=... --log.level=debug.
  • Monitor routing and escalation metrics to optimize alert workflows.