Skip to content

Production Chaos

When scaling chaos engineering to production, the goal is to balance innovation with reliability, ensuring that chaos experiments align with business objectives and system constraints. This requires integrating chaos practices into CI/CD pipelines, implementing governance frameworks, and aligning with SLOs and error budgets. Below are key strategies for transitioning chaos engineering from development to production environments.


Chaos Gates: Automating Chaos in CI/CD

Chaos gates enforce chaos experiments as mandatory checks in CI/CD pipelines, ensuring code changes are resilient before deployment. They integrate with tools like LitmusChaos or Gremlin to validate system robustness.

Example: Chaos Gate in GitLab CI

stages:
  - build
  - test
  - chaos_gate

chaos_gate:
  script:
    - litmus run --experiment=network-partition --namespace=production
  rules:
    - if: $CI_COMMIT_BRANCH == "main"
      when: always
    - if: $CI_COMMIT_BRANCH == "dev"
      when: never
This example runs a network partition experiment on the main branch, ensuring production readiness.

Diagram:

[Code Change] --> [CI/CD Pipeline] --> [Chaos Gate] --> [Deployment]


Canary Testing with Chaos: Gradual Rollouts

Canary testing involves deploying changes to a subset of users and injecting chaos to validate stability. This minimizes risk while gathering real-world feedback.

Example: Gremlin Canary Test

gremlin run --target=canary-service --chaos-type=latency --duration=300s
This command injects latency into a canary service, simulating real-world degradation scenarios.

Diagram:

[Production Traffic] --> [Canary Service]  
                  |  
                  v  
              [Chaos Injection]  


Error Budget Management: Aligning Chaos with SLOs

Error budgets quantify the "budget" of errors a system can tolerate while still meeting its SLOs. Chaos experiments should respect this budget to avoid impacting user experience.

Example: Monitoring Error Budgets with Prometheus

SELECT 
  SUM(errors) AS total_errors, 
  SUM(errors) / SUM(total_requests) * 100 AS error_rate 
FROM 
  metrics 
WHERE 
  timestamp >= now() - interval '1h'
This query calculates the error rate over the last hour, helping teams track budget usage.

Key Practice:
- Budget Thresholds: Set thresholds (e.g., 5% error rate) to trigger chaos experiments.
- Dynamic Adjustments: Scale chaos intensity based on remaining budget.


Governance and Automation: Scaling Chaos Safely

Establish governance frameworks to control chaos experiments, including approval workflows, access controls, and audit logging. Automate repetitive tasks to reduce human error.

Example: Automated Chaos Approval Workflow

if [ "$(get-error-budget-remaining)" -gt 10 ]; then  
  run-chaos-experiment --type=cpu-starvation  
else  
  echo "Error budget insufficient for chaos experiment."  
fi
This script checks the remaining error budget before initiating a chaos experiment.


Key takeaways

  • Chaos gates ensure chaos experiments are enforced in CI/CD pipelines.
  • Canary testing allows gradual validation of changes with minimal risk.
  • Error budget management ties chaos practices to SLOs and prevents over-usage.
  • Governance and automation reduce manual overhead and ensure safe, scalable chaos engineering.