Gremlin Kubernetes
Gremlin Kubernetes Chaos Setup¶
Chaos engineering on Kubernetes requires deploying Gremlin agents to monitor and execute experiments across nodes and pods. This setup enables distributed chaos scenarios, such as pod termination, network latency, and CPU throttling, while ensuring granular control over permissions and experiment execution. Below are the steps to configure Gremlin for Kubernetes chaos testing.
1. Deploy Gremlin Agents on Kubernetes¶
Gremlin agents are deployed as Kubernetes DaemonSets to ensure coverage across all nodes. The agent acts as a bridge between the Gremlin console and the cluster, enabling experiment execution and telemetry collection.
Prerequisites¶
- A running Kubernetes cluster (v1.20+).
- Access to the Gremlin console (via API or web UI).
- Cluster permissions to deploy agents and manage resources.
Deployment via Helm¶
Use the Gremlin Helm chart to install agents:
helm repo add gremlin https://charts.gremlin.com
helm install gremlin-agent gremlin/gremlin-agent \
--set agent.enabled=true \
--set console.enabled=false
nodeSelector in the Helm values.
Agent Architecture Diagram¶
2. Configure Kubernetes Permissions¶
Gremlin requires specific RBAC permissions to interact with your cluster. Create a ServiceAccount and RoleBinding to grant these permissions.
Example RBAC Configuration¶
apiVersion: v1
kind: ServiceAccount
metadata:
name: gremlin-agent
namespace: default
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: gremlin-role
rules:
- apiGroups: [""]
resources: ["pods", "nodes", "services", "endpoints"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: ["apps"]
resources: ["statefulsets", "daemonsets"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: gremlin-rolebinding
subjects:
- kind: ServiceAccount
name: gremlin-agent
namespace: default
roleRef:
kind: Role
name: gremlin-role
apiGroup: rbac.authorization.k8s
kubectl apply -f gremlin-rbac.yaml.
3. Launch Distributed Chaos Experiments¶
Once the agent is deployed and permissions are configured, use the Gremlin console to execute chaos experiments. Examples include pod termination, network latency, and CPU throttling.
Example: Pod Kill Experiment¶
curl -X POST https://console.gremlin.com/api/v1/experiments \
-H "Authorization: Bearer <API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{
"name": "PodKillExperiment",
"description": "Kill a pod in the demo namespace",
"targets": [
{
"type": "pod",
"selector": "app=demo"
}
],
"actions": [
{
"type": "kill",
"kill": {
"gracePeriodSeconds": 5
}
}
],
"duration": "60s"
}'
<API_TOKEN> with your Gremlin console API token. This experiment kills a pod with the label app=demo and runs for 60 seconds.
Monitoring Experiments¶
Use Prometheus and Grafana to monitor cluster health during experiments. Track metrics like pod restarts, latency spikes, and CPU usage to validate system resilience.
Key takeaways¶
- Deploy Gremlin agents via Helm or manifests to ensure node-level coverage.
- Configure RBAC permissions to grant the agent access to critical resources.
- Use the Gremlin console to launch targeted experiments (e.g., pod kill, network latency).
- Combine with observability tools (Prometheus/Grafana) to monitor system behavior during chaos.