Skip to content

Litmus Faults

LitmusChaos Fault Types and Use Cases

Fault injection is a cornerstone of Chaos Engineering, enabling teams to proactively test the resilience of their systems. LitmusChaos provides a rich set of fault types tailored for microservices architectures, allowing you to simulate real-world failure scenarios. Below are key fault types, their use cases, and practical examples of how to apply them.


1. Network Partition

Description: Simulates a split-brain scenario by isolating network communication between pods or services. This tests how your system handles partitioned clusters or failed network routes.
Use Case: Validate failover mechanisms in distributed databases or microservices with eventual consistency.
Example Command:

litmus run --fault network-partition \
  --namespace <namespace> \
  --selector <pod_selector> \
  --duration <duration> \
  --namespace-2 <namespace2> \
  --selector-2 <pod_selector2>
Example Scenario: Isolate two pods of a distributed database to ensure data replication continues without data loss.


2. CPU Throttling

Description: Limits CPU resources allocated to a pod, simulating resource contention or hardware degradation.
Use Case: Test how your application handles reduced computational capacity, such as during peak loads or hardware failures.
Example Command:

litmus run --fault cpu-throttling \
  --namespace <namespace> \
  --selector <pod_selector> \
  --cpu-percent <percentage> \
  --duration <duration>
Example Scenario: Apply 70% CPU throttling to a backend service to verify graceful degradation under load.


3. Disk Latency

Description: Introduces artificial delays in disk I/O operations, simulating slow storage or degraded hardware.
Use Case: Assess the impact of slow storage on read/write operations, such as in log processing pipelines.
Example Command:

litmus run --fault disk-latency \
  --namespace <namespace> \
  --selector <pod_selector> \
  --latency <milliseconds> \
  --duration <duration>
Example Scenario: Inject 500ms latency into a log aggregation service to test its ability to handle delayed data ingestion.


4. Memory Pressure

Description: Simulates memory exhaustion by allocating memory to a pod beyond its limits, triggering OOM (Out-Of-Memory) conditions.
Use Case: Validate memory management strategies and graceful shutdowns in resource-constrained environments.
Example Command:

litmus run --fault memory-pressure \
  --namespace <namespace> \
  --selector <pod_selector> \
  --memory-percent <percentage> \
  --duration <duration>
Example Scenario: Trigger memory pressure on a caching layer to ensure it evicts data properly without crashing.


5. Pod Eviction

Description: Simulates a sudden pod termination (e.g., due to node failure or resource limits), testing recovery and restart mechanisms.
Use Case: Ensure your system can handle unexpected pod failures and maintain availability.
Example Command:

litmus run --fault pod-eviction \
  --namespace <namespace> \
  --selector <pod_selector> \
  --duration <duration>
Example Scenario: Evict a frontend pod to verify that a load balancer routes traffic to healthy instances.


Applying Faults in Microservices Architectures

A typical microservices architecture includes services like APIs, databases, caches, and message brokers. Fault injection can be applied at various layers:

[Frontend Service] --> [API Gateway] --> [Microservice A] <-> [Database] <-> [Microservice B] --> [Message Broker]
        ↑                            ↓
    [Network Partition]          [Disk Latency]
  • Network Partition: Isolate the API Gateway from Microservice A to test retries or circuit breakers.
  • Disk Latency: Apply to the database to simulate slow query responses.
  • Pod Eviction: Evict Microservice B to ensure the system handles service outages.

Key takeaways

  • Network partition tests distributed system resilience and failover logic.
  • CPU throttling and disk latency help evaluate performance under resource constraints.
  • Memory pressure and pod eviction validate robustness against hardware failures.
  • Combine faults in microservices architectures to simulate complex, real-world failure scenarios.