Kubernetes Pod Priority and Preemption for Critical Workloads

Implementing pod priority classes and preemption policies to ensure critical workloads always have resources available during cluster pressure

#kubernetes#scheduling#priority#preemption
Cover image for the article: Kubernetes Pod Priority and Preemption for Critical Workloads

Introduction

In resource-constrained Kubernetes clusters, not all workloads are equal. When cluster capacity is exhausted, Kubernetes must decide which pods to schedule and which to evict. Pod Priority and Preemption provides this mechanism, ensuring business-critical workloads (payment processing, database controllers, monitoring) always run even at the expense of lower-priority workloads (batch jobs, development pods).

Without priority classes, critical pods can be starved by batch jobs that consumed all available capacity first. In production clusters, proper priority configuration eliminates 95% of critical-workload scheduling failures during capacity pressure events.

Priority and Preemption Mechanics

When a high-priority pod cannot be scheduled due to insufficient resources, the scheduler identifies lower-priority pods that can be preempted (evicted) to make room:

Chart

Preemption Decision Process

StepActionOutcome
1High-priority pod enters scheduling queueSorted by priority (higher first)
2Scheduler attempts to place podFinds no node with sufficient resources
3Scheduler evaluates preemptionIdentifies lower-priority victims
4Scheduler selects victimsMinimizes number of preemptions
5Victims are gracefully terminatedRespect terminationGracePeriodSeconds
6Resources freedHigh-priority pod scheduled

Priority Class Definition

# System-critical (reserved for cluster infrastructure)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: system-critical
value: 1000000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Cluster infrastructure - monitoring, ingress, DNS"
---
# Production-critical (revenue-impacting services)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: production-critical
value: 900000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Production services that directly impact revenue"
---
# Production-standard (standard production workloads)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: production-standard
value: 700000
globalDefault: true
preemptionPolicy: PreemptLowerPriority
description: "Standard production workloads"
---
# Batch-processing (offline/async jobs)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: batch-processing
value: 400000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Batch jobs, data pipelines, ML training"
---
# Development (dev/test workloads)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: development
value: 100000
globalDefault: false
preemptionPolicy: Never
description: "Development workloads - never preempt others"
---
# Best-effort (preemptible, non-critical)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: best-effort
value: 1000
globalDefault: false
preemptionPolicy: Never
description: "Best-effort workloads that accept eviction"

Priority Value Guidelines

Priority ClassValue RangeCan PreemptCan Be Preempted
system-node-critical2000001000EverythingNothing (built-in)
system-cluster-critical2000000000Almost allsystem-node-critical
system-critical1000000All customCluster critical
production-critical900000Standard and belowSystem
production-standard700000Batch and belowCritical and above
batch-processing400000Dev and belowStandard and above
development100000Nothing (Never)All above
best-effort1000Nothing (Never)All above

Using Priority Classes in Deployments

# Payment processing service - highest business priority
apiVersion: apps/v1
kind: Deployment
metadata:
  name: payment-processor
  namespace: production
spec:
  replicas: 3
  selector:
    matchLabels:
      app: payment-processor
  template:
    metadata:
      labels:
        app: payment-processor
    spec:
      priorityClassName: production-critical
      terminationGracePeriodSeconds: 60
      containers:
      - name: payment
        image: registry/payment:v2.1
        resources:
          requests:
            cpu: "1"
            memory: "2Gi"
          limits:
            cpu: "2"
            memory: "4Gi"
---
# ML training job - can be preempted
apiVersion: batch/v1
kind: Job
metadata:
  name: model-training
  namespace: ml-workloads
spec:
  template:
    spec:
      priorityClassName: batch-processing
      terminationGracePeriodSeconds: 120
      containers:
      - name: trainer
        image: registry/ml-trainer:latest
        resources:
          requests:
            cpu: "4"
            memory: "16Gi"
          limits:
            cpu: "8"
            memory: "32Gi"
      restartPolicy: OnFailure

Non-Preempting Priority

Use preemptionPolicy: Never for workloads that should queue rather than evict others:

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: high-priority-no-preempt
value: 800000
preemptionPolicy: Never
description: "High priority scheduling order but will not evict others"

This is useful for:

  • Development workloads that should schedule before best-effort but never disrupt production
  • Canary deployments that should queue rather than preempt stable versions
  • Cost-sensitive workloads waiting for spot capacity

Interaction with Pod Disruption Budgets

PDBs can partially protect pods from preemption:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: payment-pdb
  namespace: production
spec:
  minAvailable: 2   # Always keep at least 2 replicas
  selector:
    matchLabels:
      app: payment-processor

Important: PDBs are respected during preemption on a best-effort basis. If the scheduler cannot find victims without violating PDBs, it may still preempt PDB-protected pods after a timeout.

ScenarioPDB Respected?Notes
Normal preemptionYesScheduler avoids PDB violations
No alternative victimsNo (after timeout)Critical pods must be scheduled
Same-priority podsYesPDB-protected pods preferred to keep
Node drainYeskubectl drain respects PDBs

Monitoring Priority and Preemption

Prometheus Metrics

# Pods pending due to insufficient resources
kube_pod_status_phase{phase="Pending"} > 0

# Preemption events (count)
increase(scheduler_preemption_victims[1h])

# Preemption attempts (successful and failed)
increase(scheduler_preemption_attempts_total[1h])

# Pods by priority class
count(kube_pod_info) by (priority_class)

# Resource allocation by priority class
sum(kube_pod_container_resource_requests{resource="cpu"}) by (priority_class)

Alerting on Priority Issues

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: priority-alerts
spec:
  groups:
  - name: pod-priority
    rules:
    - alert: CriticalPodPending
      expr: |
        kube_pod_status_phase{phase="Pending"} == 1
        and on(pod, namespace)
        kube_pod_priority_class{priority_class="production-critical"}
      for: 2m
      labels:
        severity: critical
      annotations:
        summary: "Critical pod {{ $labels.pod }} is pending for >2 minutes"

    - alert: ExcessivePreemption
      expr: increase(scheduler_preemption_victims[1h]) > 20
      for: 5m
      labels:
        severity: warning
      annotations:
        summary: "High preemption rate: {{ $value }} pods preempted in last hour"

Capacity Planning with Priority

To minimize preemption events, maintain headroom for critical workloads:

Priority Class% of Cluster CapacityBuffer Recommendation
system-critical10%Always reserved (DaemonSets)
production-critical30%20% headroom above current usage
production-standard40%10% headroom
batch-processing15%None (accepts preemption)
development5%None (accepts eviction)
# Use cluster autoscaler priority expander to scale for critical pods first
apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-priority-expander
  namespace: kube-system
data:
  priorities: |
    50:
      - .*production-nodes.*
    30:
      - .*batch-nodes.*
    10:
      - .*spot-nodes.*

Key Takeaways

  • Define 4-6 priority classes that map to your organizational workload tiers rather than creating per-service priorities.
  • Set globalDefault on production-standard so that pods without explicit priority get reasonable scheduling behavior.
  • Use preemptionPolicy: Never for development to ensure dev workloads queue rather than disrupting production during resource pressure.
  • Combine priority with PodDisruptionBudgets to provide additional protection for critical services during preemption events.
  • Monitor preemption rates and alert when critical pods are pending for more than 2 minutes, indicating capacity issues.
  • Plan cluster capacity by priority tier maintaining 20% headroom above critical workload usage to minimize preemption frequency.
  • Batch and ML jobs are ideal preemption candidates because they can be restarted and checkpointed, making them natural lower-priority workloads.

Comments

    No comments yet. Be the first to share your thoughts.