Skip to content

Post-Experiment Analysis

Post-Experiment Analysis

After executing chaos experiments, thorough analysis is critical to validate system resilience, quantify impacts, and refine reliability practices. This section outlines methods for evaluating chaos experiment outcomes, including error budget consumption, SLI (Service Level Indicator) degradation, and root cause identification using logs, traces, and observability tools.


1. Quantifying Error Budget Impact

Chaos experiments can temporarily degrade system performance, consuming part of the error budget. Analyze this impact to ensure it stays within predefined thresholds.

Steps:
- Monitor metrics during and after the experiment using tools like Prometheus, Grafana, or LitmusChaos' built-in dashboards.
- Calculate error budget usage using the formula:

Error Budget Used = (Baseline SLI Violation Rate - Post-Experiment SLI Violation Rate) × Time Window  
- Example: If a service has a 0.5% error budget for 99% availability, a 1-hour experiment causing 1% errors would consume 50% of the budget.

Command to query Prometheus for error rate:

curl -G http://localhost:9090/api/v1/query \
  --data-urlencode 'query=rate(http_requests_total{status!~"2.."}[5m])' \
  --data-urlencode 'query=rate(http_requests_total{status=~"2.."}[5m])'


2. SLI Degradation Analysis

SLIs (e.g., latency, error rate, throughput) directly reflect system health. Correlate SLI spikes with chaos experiment triggers to identify vulnerabilities.

Tools:
- Grafana: Visualize SLI metrics alongside experiment timestamps.
- LitmusChaos: Use the litmusctl CLI to export experiment logs and correlate with metric anomalies.

Example:
If a latency SLI (e.g., 95th percentile) jumps from 200ms to 1500ms during a network partition experiment, this indicates a critical dependency on the disrupted component.

Command to check SLI thresholds in Grafana:

# Example: Alert if 95th percentile latency exceeds 1000ms  
curl -X POST http://localhost:3000/api/alerts -H "Content-Type: application/json" -d '{
  "rule": {
    "alert": "HighLatency",
    "expr": "percentile(http_request_duration_seconds{job=\"my-service\"}, 95) > 1000",
    "for": "5m",
    "labels": {"severity": "warning"}
  }
}'


3. Root Cause Identification with Logs and Traces

Logs and distributed tracing tools (e.g., Jaeger, Zipkin) help pinpoint where failures originated.

Steps:
- Filter logs for timestamps matching the experiment's start/end.
- Analyze traces to identify bottlenecks or failed requests.

Example:

# Search logs for errors during a chaos experiment  
grep "ERROR" /var/log/my-service.log | grep "2023-10-05T14:30:00"

Trace analysis example:
In Jaeger, trace IDs from failed requests can be used to reconstruct the flow of a request and identify which service or component failed.


4. Incident Management and Documentation

Post-experiment documentation ensures lessons are captured and repeated.

Best Practices:
- Conduct a postmortem using tools like litmusctl to export experiment details.
- Document:
- Experiment parameters (e.g., chaos type, duration).
- Observed SLI/SLO impacts.
- Root causes and mitigation steps.

Command to generate a postmortem report with LitmusChaos:

litmusctl export --experiment-name "network-partition-test" --output-format json > postmortem.json


5. Correlating Chaos Experiments with Observability Data

Use tools like Grafana Loki (logs) and Prometheus (metrics) to create dashboards that visualize:
- Chaos experiment timestamps.
- Real-time SLI/SLO metrics.
- Log entries during the experiment.

Diagram Example:

[Chaos Experiment Start] --> [Metrics Spike] --> [Log Anomalies] --> [Root Cause Identified]


Key takeaways

  • Quantify error budget usage to ensure experiments stay within acceptable thresholds.
  • Correlate SLI spikes with chaos triggers to identify critical dependencies.
  • Leverage logs and traces to isolate root causes and validate fixes.
  • Document postmortems to institutionalize learning and improve resilience.
  • Integrate observability tools to automate analysis and accelerate incident response.