Skip to content

Incident Command

Incident Command Structure

During outages, a well-defined incident command structure ensures rapid, coordinated responses while minimizing confusion and delays. This structure establishes clear roles, escalation paths, and decision-making hierarchies to maintain control and accountability. Below is a breakdown of key components and practices.


Incident Commander (IC)

The Incident Commander is the primary decision-maker during an outage, responsible for orchestrating the response and ensuring alignment with organizational goals. Key responsibilities include:
- Triage and prioritization: Assessing the severity of the incident and determining resource allocation.
- Communication: Maintaining transparency with stakeholders, including engineering teams, leadership, and external parties.
- Escalation decisions: Determining when to involve higher-level leaders or external support.
- Post-incident review: Leading the analysis to identify root causes and prevent recurrence.

Example:

# Command to trigger incident escalation (hypothetical tool)  
escalate_incident --level=2 --incident-id=INC-12345  

Diagram:

graph TD  
    A[Incident Commander] --> B[Team Leads]  
    A --> C[Stakeholders]  
    B --> D[Engineering Teams]  
    B --> E[Support Teams]  
    C --> F[Leadership]  
    C --> G[External Parties]  


Escalation Protocols

Escalation protocols define when and how incidents are escalated to higher authority or external teams. These protocols should be documented and tested regularly. Common escalation levels include:

Level Trigger Action
1 Minor outage (e.g., 5% SLI degradation) Notify on-call engineer
2 Major outage (e.g., 25% SLI degradation) Notify team lead and escalate to IC
3 Critical outage (e.g., 50% SLI degradation) Escalate to leadership and external support

Example:

# Escalation rule in a monitoring tool (e.g., Prometheus + Alertmanager)  
- alert: CriticalOutage  
  expr: (sum by (service) (job_failure_count)) > 10  
  for: 5m  
  labels:  
    severity: critical  
  annotations:  
    summary: "Critical outage detected: {{ $labels.service }}"  
    description: "Escalate to leadership and external support."  

Diagram:

graph TD  
    A[On-Call Engineer] --> B[Team Lead]  
    B --> C[Incident Commander]  
    C --> D[Leadership]  
    C --> E[External Support]  


Decision-Making Hierarchies

A clear hierarchy ensures decisions are made efficiently without bottlenecks. Key principles include:
1. Delegation: Team leads handle tactical decisions (e.g., troubleshooting), while the IC focuses on strategic decisions (e.g., resource allocation).
2. Runbooks: Standardized procedures guide decisions at each level to reduce ambiguity.
3. Authority boundaries: Engineers are empowered to make decisions within their domain, but the IC retains final authority for high-impact actions.

Example:

# Runbook snippet for resolving a database outage  
1. Verify primary DB status using `pg_isready`  
2. If unreachable, attempt failover to standby node  
3. If unresolved, escalate to DBA team (within 15 minutes)  
4. If unresolved after 30 minutes, notify IC for broader intervention  


Documentation and Post-Incident Reviews

Incident command structures must include mechanisms for documenting decisions and outcomes. This includes:
- Real-time logs: Capturing actions taken during the incident.
- Post-incident reviews (PIRs): Analyzing root causes, communication effectiveness, and lessons learned.

Example:

# Command to generate a post-incident report (hypothetical tool)  
generate_report --incident-id=INC-12345 --format=markdown  


Key takeaways

  • Clear roles (e.g., Incident Commander, team leads) prevent confusion and ensure accountability.
  • Predefined escalation paths enable rapid response without delays.
  • Decision-making hierarchies balance autonomy with oversight to avoid bottlenecks.
  • Documentation and PIRs are critical for continuous improvement and learning.