Skip to content

Performance Tuning

Performance Tuning and Scaling

Optimizing Kubernetes cluster performance requires balancing resource allocation, dynamic scaling, and efficient scheduling. This section covers techniques to fine-tune cluster behavior using resource limits, horizontal pod autoscaling (HPA), and kube-scheduler customization.


Resource Limits and Requests

Resource limits and requests define how Kubernetes allocates CPU and memory to pods. Proper configuration prevents overcommitment, ensures fair resource distribution, and avoids out-of-memory (OOM) errors.

Setting Resource Requests and Limits

Add resources.requests and resources.limits to your pod specifications. For example:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: nginx
spec:
  replicas: 3
  template:
    spec:
      containers:
      - name: nginx
        image: nginx:latest
        resources:
          requests:
            memory: "256Mi"
            cpu: "250m"
          limits:
            memory: "512Mi"
            cpu: "500m"

Best Practices

  • Requests should reflect the minimum resources required for the pod to run reliably.
  • Limits should be set based on the node's capacity and application requirements.
  • Avoid setting limits too low; this can cause pods to be evicted during high load.

Horizontal Pod Autoscaling (HPA)

HPA automatically scales the number of pod replicas based on metrics like CPU utilization or custom metrics. This ensures workloads adapt to traffic patterns without manual intervention.

Enabling HPA

Use kubectl autoscale to create an HPA:

kubectl autoscale deployment nginx --min=2 --max=10 --cpu-percent=80
This scales the nginx deployment between 2 and 10 replicas when CPU exceeds 80%.

Custom Metrics

For advanced use cases, integrate with Prometheus or the Kubernetes Metrics Server to scale based on custom metrics (e.g., request latency, queue depth). Example:

apiVersion: autoscaling/v2beta2
kind: HorizontalPodAutoscaler
metadata:
  name: nginx-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Object
    object:
      metricName: custom_metric_name
      targetValue: "100"

Limitations

  • HPA requires a metrics server (e.g., metrics-server or Prometheus).
  • Scaling decisions depend on metric granularity and update intervals.

Kube-Scheduler Tuning

The kube-scheduler assigns pods to nodes based on resource availability, affinity rules, and constraints. Customizing its behavior can improve scheduling efficiency and avoid resource bottlenecks.

Configuring Scheduler Plugins

Kube-scheduler uses plugins for scheduling and preemption. Modify the scheduler configuration to prioritize specific policies:

apiVersion: kubescheduler.config.k8s.io/v1beta3
kind: KubeSchedulerConfiguration
profiles:
- schedulerName: default-scheduler
  plugins:
    preemption:
      enabled: true
    scheduling:
      enabled: true
      name: LeastRequestedNode
This enables the LeastRequestedNode plugin, which schedules pods to nodes with the least resource demand.

Advanced Tuning

  • Use nodeSelector or affinity rules to direct workloads to specific nodes.
  • Adjust --kube-scheduler-queue-length to handle high-volume scheduling requests.
  • Monitor scheduling delays using kubectl describe pod or metrics like kube_scheduler_scheduled_pod_count.

Key takeaways

  • Resource limits and requests prevent overcommitment and ensure stable pod behavior.
  • HPA dynamically scales workloads based on metrics, improving responsiveness to traffic changes.
  • Kube-scheduler tuning optimizes node assignment, reducing bottlenecks and improving cluster efficiency.
  • Always monitor metrics and validate configurations to avoid unintended resource contention.