Skip to content

Steady State

Understanding Steady-State in Chaos

In chaos engineering, steady-state refers to the baseline operational metrics that a system maintains under normal conditions. These metrics represent the system's expected behavior and serve as a reference point to evaluate how it responds to disruptions. Steady-state metrics include latency, error rates, throughput, resource utilization (CPU, memory, disk I/O), and service availability. Establishing clear definitions for these metrics is critical to designing effective chaos experiments and validating system resilience.


Detecting Deviations During Chaos Experiments

During chaos experiments, deviations from steady-state metrics indicate potential system instability. To detect these deviations, you must:

  1. Instrument your system with monitoring tools (e.g., Prometheus, Grafana, Datadog) to collect real-time metrics.
  2. Set thresholds based on historical data and business requirements. For example, a 99th percentile latency of 200ms might be considered acceptable.
  3. Use alerts to trigger notifications when metrics exceed thresholds. For example:
    # Example Prometheus alert rule for latency spikes
    alert: HighLatency
      expr: percentiles(http_request_duration_seconds{job="my-service"}, 99) > 500
      for: 5m
      labels:
        severity: warning
    

Example: If a chaos experiment injects network latency, you might monitor the http_request_duration_seconds metric to detect if it exceeds the steady-state threshold. Tools like Grafana can visualize these metrics in real time, enabling rapid response.


Validating System Recovery

After introducing chaos, the system must return to steady-state to prove resilience. Validation involves:

  1. Automated assertions: Use scripts or tools to check if metrics return to baseline. For example:

    # Script to validate error rate returns to steady-state
    ERROR_RATE_THRESHOLD=0.01
    CURRENT_ERROR_RATE=$(curl http://localhost/metrics | grep error_rate | awk '{print $2}')
    if (( $(echo "$CURRENT_ERROR_RATE > $ERROR_RATE_THRESHOLD" | bc -l) )); then
      echo "Error rate exceeds threshold: $CURRENT_ERROR_RATE"
      exit 1
    fi
    

  2. Post-experiment analysis: Review logs, traces, and metrics to ensure the system stabilizes. Tools like distributed tracing (e.g., Jaeger) can help identify root causes of deviations.

Key consideration: Recovery validation should align with your system's error budget and SLOs. For example, if your SLO guarantees 99.9% availability, a brief dip below this threshold during chaos might be acceptable if the system recovers within the error budget.


Diagram: Steady-State and Chaos Orchestration

+-------------------+       +-------------------+       +-------------------+
|   Steady-State    | ----> | Chaos Experiment | ----> |   Recovery Check   |
| (Latency, Error   |       | (Network Latency, |       | (Metrics Thresholds)|
| Rate, Throughput) |       | etc.)            |       +-------------------+
+-------------------+       +-------------------+       |
          ^                           ^                       |
          |                           |                       |
          v                           v                       v
+-------------------+       +-------------------+       +-------------------+
|  Monitoring Tools | <---- |  Alerting System  | <---- |  Postmortem Tools  |
| (Prometheus,      |       | (Grafana,         |       | (Chaos Monkey,    |
| Grafana)          |       | Loki)            |       | etc.)             |
+-------------------+       +-------------------+       +-------------------+

Key takeaways

  • Steady-state metrics are the baseline for system behavior; define them clearly using historical data and business requirements.
  • Real-time monitoring with tools like Prometheus and Grafana enables rapid detection of deviations during chaos experiments.
  • Automated validation ensures the system returns to steady-state after chaos, aligning with SLOs and error budgets.
  • Post-experiment analysis is critical to understanding root causes and improving resilience.