Skip to content

Advanced Scenarios

Advanced Kubernetes Chaos Scenarios

Kubernetes-native applications require rigorous testing of resilience under complex, real-world failure conditions. Advanced chaos scenarios simulate multi-dimensional failures that challenge system robustness, such as pod eviction, storage I/O latency, and network latency injection. These experiments help validate fault tolerance, auto-scaling, and recovery mechanisms in distributed systems.


Pod Eviction Scenario

Overview
Pod eviction simulates node resource exhaustion (CPU, memory, or disk) to trigger Kubernetes' eviction mechanisms. This tests how applications handle sudden pod termination and whether recovery processes (e.g., restarts, rescheduling) are effective.

Example Command

gremlin attack run k8s-pod-eviction \
  --namespace <namespace> \
  --pod-name <pod-name> \
  --reason "ResourceExhaustion" \
  --duration 30s
This command evicts a specific pod by triggering a ResourceExhaustion event, forcing Kubernetes to terminate the pod.

Diagram Description
A diagram would show:
1. A Kubernetes node running a pod.
2. Gremlin injecting a resource exhaustion event.
3. The kubelet evicting the pod.
4. The kube-scheduler rescheduling the pod to another node.

Key Considerations
- Ensure the pod has a restartPolicy of Always or OnFailure for meaningful results.
- Monitor metrics like kube_pod_eviction_total in Prometheus to track eviction events.


Storage I/O Latency Injection

Overview
Storage I/O latency simulates delays in read/write operations to persistent volumes (PVs), testing how applications handle degraded storage performance. This is critical for databases, stateful apps, or workloads relying on fast disk access.

Example Command

gremlin attack run k8s-storage-io-latency \
  --namespace <namespace> \
  --storage-class <storage-class> \
  --latency 500ms \
  --duration 60s
This command introduces a 500ms delay on all I/O operations for the specified storage class, mimicking slow disk access.

Diagram Description
A diagram would show:
1. A pod accessing a PV.
2. Gremlin injecting latency into the storage I/O path.
3. The application experiencing increased query latency or timeouts.
4. The system's response (e.g., retries, fallback mechanisms).

Key Considerations
- Test with storage classes used by critical workloads.
- Use Prometheus metrics like node_disk_read_time_seconds to monitor I/O performance.


Network Latency Injection

Overview
Network latency injection simulates slow or unstable network connections between services, testing microservices resilience. This is vital for distributed systems where communication delays can cascade into failures.

Example Command

gremlin attack run k8s-network-latency \
  --namespace <namespace> \
  --target <service-name> \
  --latency 1500ms \
  --protocol tcp \
  --duration 45s
This command adds 1500ms latency to all TCP traffic to the specified service, simulating a high-latency network link.

Diagram Description
A diagram would show:
1. Two pods/services communicating over a network.
2. Gremlin injecting latency into the network path.
3. The receiving service experiencing delayed responses.
4. The system's handling of timeouts or retries.

Key Considerations
- Use tools like tcpdump or Wireshark to validate latency injection.
- Monitor service-level metrics (e.g., http_request_duration_seconds) for performance degradation.


Key takeaways

  • Pod eviction tests resilience to sudden node failures and rescheduling.
  • Storage I/O latency validates stateful workloads under degraded disk performance.
  • Network latency ensures microservices handle slow or unstable communication.
  • Combine these scenarios with observability tools (Prometheus, Grafana) to analyze system behavior.
  • Always validate chaos experiments with rollback strategies and postmortem analysis.