Skip to content

LitmusChaos Best Practices

Chaos engineering with LitmusChaos requires careful planning to ensure experiments are meaningful, repeatable, and aligned with your system's reliability goals. This section outlines best practices for designing experiments with measurable SLI metrics, avoiding false positives, and integrating chaos testing into CI/CD pipelines.


Designing Experiments with Measurable SLI Metrics

Align Fault Injection with SLIs/SLOs

Every chaos experiment should directly tie to specific SLIs (e.g., latency, error rate, availability) and SLOs. For example, if your SLO guarantees 99.9% availability, design experiments to simulate outages and measure how the system responds.
Example:

apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: nginx-chaos
spec:
  engineName: nginx-chaos
  annotation:
    chaos.io/namespace: default
  workloadSelector:
    workloadType: nginx
    workloadLabel: app=nginx
  chaosServiceAccount: litmus-chaos-sa
  experiments:
  - name: network-latency
    spec:
      components:
        - component: nginx
          metrics:
            - name: latency
              threshold: "500"
              interval: "10s"
              severity: "critical"

Use Monitoring Tools for Real-Time Feedback

Integrate Prometheus/Grafana to visualize SLI metrics during experiments. For example, a Prometheus query to track HTTP error rates:

sum by (job) (rate(http_requests_total{job="nginx", status!~"2.."}[5m]))

Example: Chaos Experiment with SLI Validation

kubectl apply -f chaos-experiment.yaml
# Monitor metrics in Grafana during the experiment

Avoiding False Positives

Use Canary Releases and Gradual Injection

Start with low-intensity faults (e.g., 10% network latency) and gradually increase severity. Deploy chaos experiments to a canary subset of your workload to minimize risk.
Example:

spec:
  components:
    - component: nginx
      parameters:
        - name: delay
          value: "10s"
      # Gradually increase delay in subsequent runs

Validate System Resilience with LitmusChaos

Leverage LitmusChaos' built-in validation checks to ensure the system recovers and metrics remain within acceptable bounds.
Example:

spec:
  validation:
    - metric: latency
      threshold: "500"
      interval: "10s"
      failure: true

Example: Post-Experiment Validation

# Check if all validations passed
kubectl get chaosengine nginx-chaos -o jsonpath='{.status.phase}'
# Output: "Completed" if all validations succeeded

Integrating with CI/CD Pipelines

Automate Chaos Tests in CI/CD

Embed chaos experiments into your pipeline to catch regressions early. For example, run chaos tests in a staging environment before merging to production.
Example GitHub Actions Workflow:

jobs:
  chaos-test:
    runs-on: ubuntu-latest
    steps:
      - name: Deploy staging environment
        run: kubectl apply -f staging-deployment.yaml
      - name: Run chaos experiment
        run: kubectl apply -f chaos-experiment.yaml
      - name: Validate results
        run: |
          if [ "$(kubectl get chaosengine nginx-chaos -o jsonpath='{.status.phase}')" == "Completed" ]; then
            echo "Chaos test passed"
          else
            echo "Chaos test failed"
            exit 1
          fi

Handle Failures Gracefully

If an experiment fails, pause the pipeline and investigate. Use postmortems to analyze root causes and improve resilience.
Example:

# Fail the pipeline if chaos test fails
if [ "$CHAOSENGINE_PHASE" != "Completed" ]; then
  echo "Chaos test failed: $CHAOSENGINE_PHASE"
  exit 1
fi


Key takeaways

  • Align experiments with SLIs/SLOs to ensure measurable outcomes.
  • Use validation checks to avoid false positives and verify system resilience.
  • Automate chaos testing in CI/CD pipelines to catch regressions early.
  • Monitor metrics in real time using tools like Prometheus and Grafana.
  • Prioritize gradual fault injection and canary releases to minimize risk.