Skip to content

Game Days

Planning and Executing Game Days

Game Days are coordinated chaos experiments designed to simulate real-world failure scenarios across systems, teams, and infrastructure. These events help SRE teams validate resilience, improve incident response, and align cross-functional collaboration. Success requires meticulous planning, realistic scenario design, and rigorous post-experiment analysis.


1. Planning Phase: Objectives and Team Alignment

Define Clear Goals

  • Align Game Days with SLOs and error budgets to identify critical systems for testing. For example, target services with high SLI thresholds or mission-critical dependencies.
  • Example: "Simulate a regional AWS outage to test failover mechanisms for our global API gateway."

Assemble a Cross-Team Team

  • Include developers, SREs, DevOps, and business stakeholders to ensure diverse perspectives.
  • Assign roles: Chaos Engineer (orchestrates experiments), Observer (monitors metrics), Incident Commander (manages real-time response), and Documenter (records findings).

Select Tools and Scenarios

  • Use LitmusChaos or Gremlin to inject failures (e.g., network latency, CPU throttling, database outages).
  • Design scenarios that mimic real-world incidents, such as:
  • Simulating a database replica failure.
  • Introducing latency in a critical API endpoint.
  • Corrupting a config file in a distributed system.

Create a Communication Plan

  • Define escalation paths, on-call rotation schedules, and pre-approved rollback procedures.
  • Use tools like Slack or Microsoft Teams for real-time updates.

2. Execution Phase: Running the Experiment

Set Up the Environment

  • Isolate the experiment to avoid unintended impacts. For example, use staging environments or canary deployments.
  • Example command to start a chaos experiment with LitmusChaos:
    kubectl apply -f https://raw.githubusercontent.com/litmuschaos/litmus-chaos/master/charts/litmuschaos/crds/litmuschaos_cr.yaml
    

Run Simulated Failures

  • Use Gremlin to inject chaos:
    gremlin create --type latency --target "http://api-gateway:8080" --duration 300 --latency 500
    
  • Monitor system behavior using Prometheus and Grafana dashboards. Track SLI metrics (e.g., request latency, error rates) in real time.

Simulate Coordination Challenges

  • Introduce correlated failures (e.g., a network partition followed by a database crash) to test system resilience under complex conditions.
  • Example: Simulate a DNS outage followed by a load balancer misconfiguration to test failover workflows.

3. Post-Execution: Analysis and Learning

Analyze System Behavior

  • Compare pre- and post-chaos metrics to assess impact. Use Grafana to visualize anomalies (e.g., spike in error rates, latency increases).
  • Example:
    SELECT * FROM metrics WHERE metric_name = 'http_request_latency' AND timestamp > '2023-10-01T00:00:00Z'
    

Document Findings

  • Record:
  • Which failures caused cascading issues.
  • How teams responded to incidents.
  • Gaps in monitoring or automation.
  • Use tools like Confluence or Notion to centralize documentation.

Conduct a Postmortem

  • Follow the 5 Whys framework to root-cause failures:
  • Why did the system fail?
  • Why was the failure undetected?
  • Why wasn’t the rollback plan executed?
  • Why were dependencies not considered?
  • What can be done to prevent recurrence?

4. Iterate and Improve

  • Update SLOs or error budgets based on lessons learned.
  • Automate repetitive tasks (e.g., rollback procedures) using Ansible or Terraform.
  • Schedule regular Game Days to maintain system resilience.

Key takeaways

  • Cross-team collaboration is critical for realistic chaos scenarios.
  • Realistic failure chains (e.g., network + database failures) reveal hidden dependencies.
  • Observability tools (Prometheus, Grafana) are essential for real-time analysis.
  • Postmortems should focus on systemic improvements, not just blame.
  • Iterate based on findings to strengthen system resilience over time.