Kubernetes Pod Priority and Preemption for Critical Workloads
Implementing pod priority classes and preemption policies to ensure critical workloads always have resources available during cluster pressure

Introduction
In resource-constrained Kubernetes clusters, not all workloads are equal. When cluster capacity is exhausted, Kubernetes must decide which pods to schedule and which to evict. Pod Priority and Preemption provides this mechanism, ensuring business-critical workloads (payment processing, database controllers, monitoring) always run even at the expense of lower-priority workloads (batch jobs, development pods).
Without priority classes, critical pods can be starved by batch jobs that consumed all available capacity first. In production clusters, proper priority configuration eliminates 95% of critical-workload scheduling failures during capacity pressure events.
Priority and Preemption Mechanics
When a high-priority pod cannot be scheduled due to insufficient resources, the scheduler identifies lower-priority pods that can be preempted (evicted) to make room:
Preemption Decision Process
| Step | Action | Outcome |
|---|---|---|
| 1 | High-priority pod enters scheduling queue | Sorted by priority (higher first) |
| 2 | Scheduler attempts to place pod | Finds no node with sufficient resources |
| 3 | Scheduler evaluates preemption | Identifies lower-priority victims |
| 4 | Scheduler selects victims | Minimizes number of preemptions |
| 5 | Victims are gracefully terminated | Respect terminationGracePeriodSeconds |
| 6 | Resources freed | High-priority pod scheduled |
Priority Class Definition
Recommended Priority Hierarchy
# System-critical (reserved for cluster infrastructure)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: system-critical
value: 1000000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Cluster infrastructure - monitoring, ingress, DNS"
---
# Production-critical (revenue-impacting services)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: production-critical
value: 900000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Production services that directly impact revenue"
---
# Production-standard (standard production workloads)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: production-standard
value: 700000
globalDefault: true
preemptionPolicy: PreemptLowerPriority
description: "Standard production workloads"
---
# Batch-processing (offline/async jobs)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-processing
value: 400000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Batch jobs, data pipelines, ML training"
---
# Development (dev/test workloads)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: development
value: 100000
globalDefault: false
preemptionPolicy: Never
description: "Development workloads - never preempt others"
---
# Best-effort (preemptible, non-critical)
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: best-effort
value: 1000
globalDefault: false
preemptionPolicy: Never
description: "Best-effort workloads that accept eviction"
Priority Value Guidelines
| Priority Class | Value Range | Can Preempt | Can Be Preempted |
|---|---|---|---|
| system-node-critical | 2000001000 | Everything | Nothing (built-in) |
| system-cluster-critical | 2000000000 | Almost all | system-node-critical |
| system-critical | 1000000 | All custom | Cluster critical |
| production-critical | 900000 | Standard and below | System |
| production-standard | 700000 | Batch and below | Critical and above |
| batch-processing | 400000 | Dev and below | Standard and above |
| development | 100000 | Nothing (Never) | All above |
| best-effort | 1000 | Nothing (Never) | All above |
Using Priority Classes in Deployments
# Payment processing service - highest business priority
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-processor
namespace: production
spec:
replicas: 3
selector:
matchLabels:
app: payment-processor
template:
metadata:
labels:
app: payment-processor
spec:
priorityClassName: production-critical
terminationGracePeriodSeconds: 60
containers:
- name: payment
image: registry/payment:v2.1
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
---
# ML training job - can be preempted
apiVersion: batch/v1
kind: Job
metadata:
name: model-training
namespace: ml-workloads
spec:
template:
spec:
priorityClassName: batch-processing
terminationGracePeriodSeconds: 120
containers:
- name: trainer
image: registry/ml-trainer:latest
resources:
requests:
cpu: "4"
memory: "16Gi"
limits:
cpu: "8"
memory: "32Gi"
restartPolicy: OnFailure
Non-Preempting Priority
Use preemptionPolicy: Never for workloads that should queue rather than evict others:
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: high-priority-no-preempt
value: 800000
preemptionPolicy: Never
description: "High priority scheduling order but will not evict others"
This is useful for:
- Development workloads that should schedule before best-effort but never disrupt production
- Canary deployments that should queue rather than preempt stable versions
- Cost-sensitive workloads waiting for spot capacity
Interaction with Pod Disruption Budgets
PDBs can partially protect pods from preemption:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payment-pdb
namespace: production
spec:
minAvailable: 2 # Always keep at least 2 replicas
selector:
matchLabels:
app: payment-processor
Important: PDBs are respected during preemption on a best-effort basis. If the scheduler cannot find victims without violating PDBs, it may still preempt PDB-protected pods after a timeout.
| Scenario | PDB Respected? | Notes |
|---|---|---|
| Normal preemption | Yes | Scheduler avoids PDB violations |
| No alternative victims | No (after timeout) | Critical pods must be scheduled |
| Same-priority pods | Yes | PDB-protected pods preferred to keep |
| Node drain | Yes | kubectl drain respects PDBs |
Monitoring Priority and Preemption
Prometheus Metrics
# Pods pending due to insufficient resources
kube_pod_status_phase{phase="Pending"} > 0
# Preemption events (count)
increase(scheduler_preemption_victims[1h])
# Preemption attempts (successful and failed)
increase(scheduler_preemption_attempts_total[1h])
# Pods by priority class
count(kube_pod_info) by (priority_class)
# Resource allocation by priority class
sum(kube_pod_container_resource_requests{resource="cpu"}) by (priority_class)
Alerting on Priority Issues
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: priority-alerts
spec:
groups:
- name: pod-priority
rules:
- alert: CriticalPodPending
expr: |
kube_pod_status_phase{phase="Pending"} == 1
and on(pod, namespace)
kube_pod_priority_class{priority_class="production-critical"}
for: 2m
labels:
severity: critical
annotations:
summary: "Critical pod {{ $labels.pod }} is pending for >2 minutes"
- alert: ExcessivePreemption
expr: increase(scheduler_preemption_victims[1h]) > 20
for: 5m
labels:
severity: warning
annotations:
summary: "High preemption rate: {{ $value }} pods preempted in last hour"
Capacity Planning with Priority
To minimize preemption events, maintain headroom for critical workloads:
| Priority Class | % of Cluster Capacity | Buffer Recommendation |
|---|---|---|
| system-critical | 10% | Always reserved (DaemonSets) |
| production-critical | 30% | 20% headroom above current usage |
| production-standard | 40% | 10% headroom |
| batch-processing | 15% | None (accepts preemption) |
| development | 5% | None (accepts eviction) |
# Use cluster autoscaler priority expander to scale for critical pods first
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-autoscaler-priority-expander
namespace: kube-system
data:
priorities: |
50:
- .*production-nodes.*
30:
- .*batch-nodes.*
10:
- .*spot-nodes.*
Key Takeaways
- Define 4-6 priority classes that map to your organizational workload tiers rather than creating per-service priorities.
- Set globalDefault on production-standard so that pods without explicit priority get reasonable scheduling behavior.
- Use preemptionPolicy: Never for development to ensure dev workloads queue rather than disrupting production during resource pressure.
- Combine priority with PodDisruptionBudgets to provide additional protection for critical services during preemption events.
- Monitor preemption rates and alert when critical pods are pending for more than 2 minutes, indicating capacity issues.
- Plan cluster capacity by priority tier maintaining 20% headroom above critical workload usage to minimize preemption frequency.
- Batch and ML jobs are ideal preemption candidates because they can be restarted and checkpointed, making them natural lower-priority workloads.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.