Skip to content

Validating Fencing

Testing and Validating Fencing is a critical step in ensuring the reliability of a Pacemaker cluster. Fencing mechanisms (STONITH) must be verified to ensure they can isolate failed nodes and prevent data corruption or split-brain scenarios. This section outlines methods to simulate node failures, validate fencing behavior, and confirm the effectiveness of your STONITH configuration.


Simulating Node Failures

To test fencing, simulate node failures without disrupting production workloads. Use the following approaches:

1. Use pcs cluster simulate

Simulate a node failure by triggering a simulated failure scenario:

pcs cluster simulate <node>
Monitor the cluster response using crm_mon -l to observe if the cluster detects the failure and initiates fencing. For virtual machines, you can also simulate a crash by:
pcs cluster simulate <node>
Ensure the cluster isolates the node and transitions it to a fenced state.

2. Test STONITH Devices Directly

Use the stonith_admin utility to test fencing devices:

sudo stonith_admin -t <stonith_device_name>
Replace <stonith_device_name> with your actual fencing device (e.g., fence_xvm for VMs). This command verifies the device's ability to power off or isolate a node.


Verifying Fencing Effectiveness

After simulating failures, validate that the cluster correctly isolates nodes and maintains quorum:

1. Check Cluster Status

Use crm_mon -l to confirm the cluster's state. A properly fenced node should appear in a fenced state, and the cluster should retain quorum:

Node 1: online
Node 2: fenced

2. Inspect STONITH Device Status

Verify the STONITH device's operational status:

pcs stonith status
Ensure the device reports online and has no pending errors.

3. Review Logs

Check logs for fencing-related events:

journalctl -u pacemaker --since "1 hour ago"
Look for entries like stonith-action or fence_xvm to confirm fencing actions were executed.


Best Practices for Fencing Testing

  • Avoid Production Systems: Test in a staging environment or isolated VMs to prevent unintended outages.
  • Document Scenarios: Record test cases and results to identify edge cases or configuration gaps.
  • Validate Device Compatibility: Ensure fencing devices are supported by your cluster version (e.g., IPMI, power management tools).
  • Test Recovery: After fencing, verify the node can rejoin the cluster by restarting its services:
    sudo systemctl start pacemaker
    

Key takeaways

  • Simulate node failures using pcs cluster simulate or direct STONITH testing to validate fencing behavior.
  • Monitor cluster status with crm_mon and check STONITH device logs to confirm isolation success.
  • Always test in non-production environments to avoid disruptions.
  • Regularly validate fencing configurations to ensure resilience against hardware or software failures.
  • Document test outcomes to refine STONITH policies and improve cluster reliability.