Litmus Faults
LitmusChaos Fault Types and Use Cases¶
Fault injection is a cornerstone of Chaos Engineering, enabling teams to proactively test the resilience of their systems. LitmusChaos provides a rich set of fault types tailored for microservices architectures, allowing you to simulate real-world failure scenarios. Below are key fault types, their use cases, and practical examples of how to apply them.
1. Network Partition¶
Description: Simulates a split-brain scenario by isolating network communication between pods or services. This tests how your system handles partitioned clusters or failed network routes.
Use Case: Validate failover mechanisms in distributed databases or microservices with eventual consistency.
Example Command:
litmus run --fault network-partition \
--namespace <namespace> \
--selector <pod_selector> \
--duration <duration> \
--namespace-2 <namespace2> \
--selector-2 <pod_selector2>
2. CPU Throttling¶
Description: Limits CPU resources allocated to a pod, simulating resource contention or hardware degradation.
Use Case: Test how your application handles reduced computational capacity, such as during peak loads or hardware failures.
Example Command:
litmus run --fault cpu-throttling \
--namespace <namespace> \
--selector <pod_selector> \
--cpu-percent <percentage> \
--duration <duration>
3. Disk Latency¶
Description: Introduces artificial delays in disk I/O operations, simulating slow storage or degraded hardware.
Use Case: Assess the impact of slow storage on read/write operations, such as in log processing pipelines.
Example Command:
litmus run --fault disk-latency \
--namespace <namespace> \
--selector <pod_selector> \
--latency <milliseconds> \
--duration <duration>
4. Memory Pressure¶
Description: Simulates memory exhaustion by allocating memory to a pod beyond its limits, triggering OOM (Out-Of-Memory) conditions.
Use Case: Validate memory management strategies and graceful shutdowns in resource-constrained environments.
Example Command:
litmus run --fault memory-pressure \
--namespace <namespace> \
--selector <pod_selector> \
--memory-percent <percentage> \
--duration <duration>
5. Pod Eviction¶
Description: Simulates a sudden pod termination (e.g., due to node failure or resource limits), testing recovery and restart mechanisms.
Use Case: Ensure your system can handle unexpected pod failures and maintain availability.
Example Command:
litmus run --fault pod-eviction \
--namespace <namespace> \
--selector <pod_selector> \
--duration <duration>
Applying Faults in Microservices Architectures¶
A typical microservices architecture includes services like APIs, databases, caches, and message brokers. Fault injection can be applied at various layers:
[Frontend Service] --> [API Gateway] --> [Microservice A] <-> [Database] <-> [Microservice B] --> [Message Broker]
↑ ↓
[Network Partition] [Disk Latency]
- Network Partition: Isolate the API Gateway from Microservice A to test retries or circuit breakers.
- Disk Latency: Apply to the database to simulate slow query responses.
- Pod Eviction: Evict Microservice B to ensure the system handles service outages.
Key takeaways¶
- Network partition tests distributed system resilience and failover logic.
- CPU throttling and disk latency help evaluate performance under resource constraints.
- Memory pressure and pod eviction validate robustness against hardware failures.
- Combine faults in microservices architectures to simulate complex, real-world failure scenarios.