Cluster Alerts
Pacemaker clusters rely on timely alerts to notify administrators of critical events such as node failures, resource downtime, or configuration issues. This section covers methods to configure alerts using systemd, email, and external monitoring tools, ensuring proactive maintenance and rapid response.
Using systemd for Cluster Alerts¶
Systemd can be leveraged to run custom scripts that monitor cluster health and trigger alerts. Create a systemd service unit to periodically check the cluster status and execute alert actions.
Example: Create a systemd service for cluster health checks
[Unit]
Description=Cluster Health Alert Service
Timer=cluster-alert.timer
[Service]
Type=simple
ExecStart=/usr/bin/cluster-health-check.sh
Example: Define a timer to run checks every 5 minutes
[Unit]
Description=Run cluster health checks every 5 minutes
[Timer]
OnCalendar=*:0/5
Unit=cluster-alert.service
Example: Script to trigger alerts
#!/bin/bash
CLUSTER_STATUS=$(crm_mon -1 | grep "stack is online")
if [[ ! $CLUSTER_STATUS ]]; then
echo "Cluster is offline!" | mail -s "Pacemaker Alert" [email protected]
fi
Email Alerts with mail or sendmail¶
Configure the cluster to send email notifications for critical events. Use tools like mail or sendmail to relay alerts.
Example: Script to send email alerts
#!/bin/bash
if ! crm_mon -1 | grep -q "online"; then
echo "CRITICAL: Cluster is offline!" | mail -s "Pacemaker Alert" [email protected]
fi
postfix or sendmail). Note that the mail command requires an MTA to be installed and properly configured before use.
External Monitoring Tools¶
Integrate Pacemaker with external tools like Nagios, Zabbix, or Prometheus for centralized monitoring. Use the Pacemaker REST API or SSH to query cluster status.
Example: Check cluster status via SSH
ssh cluster-node1 "crm_mon -1 | grep 'online'"
if [[ $? -ne 0 ]]; then
echo "Cluster offline on cluster-node1" | tee /path/to/alert.log
fi
Example: Nagios plugin for cluster health
#!/bin/bash
if crm_mon -1 | grep -q "online"; then
echo "OK"
exit 0
else
echo "CRITICAL: Cluster offline"
exit 2
fi
nagios.cfg.
Testing and Validation¶
- Simulate node failure:
Note: This command requires a multi-node cluster. For single-node setups, use alternative methods like stopping the cluster service or simulating failures through other means. - Verify alerts are triggered via email, logs, or monitoring tools.
- Restore the node and ensure alerts resolve.
Key takeaways¶
- Use systemd to run periodic health checks and trigger alerts via scripts.
- Configure email alerts using
mailorsendmailfor quick notifications. - Integrate with external tools like Nagios or Zabbix for centralized monitoring.
- Test alerts by simulating failures to ensure reliability.
- Always validate alert configurations against your environment's specific tools and policies.