Zero-Downtime Kubernetes Rolling Updates: Pod Disruption Budgets and Graceful Termination

Achieving true zero-downtime deployments in Kubernetes with properly configured PDBs, preStop hooks, readiness gates, and graceful shutdown patterns.

#kubernetes#deployment#rolling-update#high-availability
Cover image for the article: Zero-Downtime Kubernetes Rolling Updates: Pod Disruption Budgets and Graceful Termination

"Zero downtime" in Kubernetes is a promise that requires careful configuration to fulfill. The default rolling update strategy replaces pods, but without proper graceful termination, readiness signaling, and disruption budgets, users experience dropped connections, failed requests, and intermittent errors during every deployment. True zero-downtime requires coordination between the application, Kubernetes, load balancers, and the deployment strategy.

After instrumenting 94 production services with request-level deployment impact tracking, we achieved verified zero-downtime across 18,000+ deployments. Here is every configuration that matters.

The Problem: "Rolling Update" Is Not Zero-Downtime by Default

A standard Kubernetes Deployment with rolling update strategy has multiple failure windows:

  • In-flight requests dropped: Pod receives SIGTERM while processing requests. Default: immediate termination.
  • Load balancer race condition: Pod removed from service endpoints, but load balancer still routes traffic to it for up to 30 seconds.
  • Readiness gap: New pods accept traffic before the application is fully warmed up (connection pools, caches, JIT compilation).
  • Thundering herd: All pods updating simultaneously under maxUnavailable: 25% can overwhelm remaining capacity.

Our measurement showed that a "standard" rolling update caused 0.3-0.8% request failures during a 3-minute deployment window — unacceptable for services handling financial transactions.

Architecture: The Zero-Downtime Stack

True zero-downtime requires coordination at four layers:

Zero-Downtime Deployment Stack

Layer 1: Graceful Shutdown with preStop Hook

The pod must continue serving in-flight requests after receiving SIGTERM. A preStop hook delays the shutdown to allow the load balancer to drain connections:

# deployments/api-gateway/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-gateway
spec:
  replicas: 6
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 2
      maxUnavailable: 0  # Never reduce available capacity
  template:
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: api-gateway
          image: ecr.aws/org/api-gateway:v3.2.1
          ports:
            - containerPort: 8080
          lifecycle:
            preStop:
              exec:
                command:
                  - /bin/sh
                  - -c
                  - |
                    # Signal the app to stop accepting new connections
                    kill -SIGTERM 1
                    # Wait for load balancer to deregister (AWS ALB: ~15s)
                    sleep 15
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 5
            failureThreshold: 2
          livenessProbe:
            httpGet:
              path: /health/live
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 10
            failureThreshold: 3
          resources:
            requests:
              memory: "512Mi"
              cpu: "500m"
            limits:
              memory: "1Gi"
              cpu: "1000m"

Why maxUnavailable: 0?

Setting maxUnavailable: 0 ensures Kubernetes never removes a pod until its replacement is ready. Combined with maxSurge: 2, the cluster temporarily runs extra capacity during updates rather than reducing capacity.

Layer 2: Application-Level Graceful Shutdown

The application must handle SIGTERM by stopping new request acceptance while completing in-flight work:

// src/server.ts
import { createServer, Server } from 'http';
import { promisify } from 'util';

class GracefulServer {
  private server: Server;
  private isShuttingDown = false;
  private activeConnections = new Set<any>();
  private metrics: PrometheusMetrics;

  constructor(private app: Express) {
    this.server = createServer(app);
    this.metrics = new PrometheusMetrics();

    // Track active connections
    this.server.on('connection', (conn) => {
      this.activeConnections.add(conn);
      conn.on('close', () => this.activeConnections.delete(conn));
    });

    // Middleware to reject new requests during shutdown
    this.app.use((req, res, next) => {
      if (this.isShuttingDown) {
        res.set('Connection', 'close');
        res.status(503).json({
          error: 'Service shutting down',
          retryAfter: 5
        });
        return;
      }
      next();
    });
  }

  async start(port: number): Promise<void> {
    this.server.listen(port, () => {
      console.log(`Server listening on port ${port}`);
    });

    // Handle shutdown signals
    process.on('SIGTERM', () => this.shutdown('SIGTERM'));
    process.on('SIGINT', () => this.shutdown('SIGINT'));
  }

  private async shutdown(signal: string): Promise<void> {
    console.log(`Received ${signal}. Starting graceful shutdown...`);
    this.isShuttingDown = true;
    this.metrics.increment('graceful_shutdown_initiated');

    // Stop accepting new connections
    const closeServer = promisify(this.server.close.bind(this.server));

    // Wait for active requests to complete (max 30s)
    const shutdownTimeout = setTimeout(() => {
      console.warn('Shutdown timeout reached. Forcing close.');
      this.activeConnections.forEach(conn => conn.destroy());
      this.metrics.increment('graceful_shutdown_timeout');
    }, 30000);

    try {
      await closeServer();
      clearTimeout(shutdownTimeout);
      console.log(`Graceful shutdown complete. ${this.activeConnections.size} connections closed.`);
      this.metrics.increment('graceful_shutdown_success');
    } catch (err) {
      console.error('Error during shutdown:', err);
      this.metrics.increment('graceful_shutdown_error');
    } finally {
      process.exit(0);
    }
  }
}

Layer 3: Pod Disruption Budgets

PDBs prevent Kubernetes from evicting too many pods simultaneously during voluntary disruptions (node drains, cluster upgrades, spot instance reclamation):

# deployments/api-gateway/pdb.yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-gateway-pdb
spec:
  # At least 4 of 6 pods must remain available during disruptions
  minAvailable: 4
  selector:
    matchLabels:
      app: api-gateway
---
# For critical services: percentage-based PDB
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: payment-processor-pdb
spec:
  # Never allow more than 1 pod unavailable regardless of replica count
  maxUnavailable: 1
  selector:
    matchLabels:
      app: payment-processor

Layer 4: Readiness Gates for Load Balancer Integration

AWS Load Balancer Controller uses readiness gates to ensure the ALB target is registered and healthy before the pod receives traffic:

# deployments/api-gateway/target-group-binding.yaml
apiVersion: elbv2.k8s.aws/v1beta1
kind: TargetGroupBinding
metadata:
  name: api-gateway-tgb
spec:
  serviceRef:
    name: api-gateway
    port: 8080
  targetGroupARN: arn:aws:elasticloadbalancing:us-east-1:123456789:targetgroup/api-gw/abc123
  targetType: ip
  # Pods won't receive traffic until ALB health check passes
  networking:
    ingress:
      - from:
          - securityGroup:
              groupID: sg-alb-security-group
        ports:
          - port: 8080
            protocol: TCP

The pod's readiness condition includes target-health.elbv2.k8s.aws/api-gateway-tgb, ensuring it is not marked ready until the ALB confirms it healthy.

Startup Probes for Slow-Starting Services

Services with initialization work (loading ML models, warming caches, establishing connection pools) need startup probes to prevent premature traffic:

# For services that take 30-60s to initialize
startupProbe:
  httpGet:
    path: /health/startup
    port: 8080
  initialDelaySeconds: 10
  periodSeconds: 5
  failureThreshold: 12  # Up to 70s for startup
  successThreshold: 1

The startup probe runs before readiness and liveness probes activate. This prevents slow-starting pods from being killed (liveness) or receiving traffic (readiness) prematurely.

Measuring Zero-Downtime: Deployment Impact Score

We track a Deployment Impact Score (DIS) — the percentage of requests that fail specifically during deployment windows:

# monitoring/deployment-impact.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: deployment-impact-score
spec:
  groups:
    - name: deployment-impact
      rules:
        - record: deployment:impact_score:5m
          expr: |
            (
              sum(rate(http_requests_total{code=~"5.."}[5m])
                * on(namespace, deployment) group_left()
                kube_deployment_status_observed_generation
                != on(namespace, deployment)
                kube_deployment_metadata_generation)
              /
              sum(rate(http_requests_total[5m])
                * on(namespace, deployment) group_left()
                kube_deployment_status_observed_generation
                != on(namespace, deployment)
                kube_deployment_metadata_generation)
            ) * 100

        - alert: DeploymentImpactDetected
          expr: deployment:impact_score:5m > 0.01
          for: 1m
          labels:
            severity: warning
          annotations:
            summary: "Deployment causing request failures: {{ $value }}% error rate"

Production Results

After implementing the full zero-downtime stack across 94 services:

Deployment Impact Results

MetricBeforeAfterImprovement
Request failures during deploy0.3-0.8%0.000%Eliminated
Connection resets during deploy12/deploy0/deployEliminated
Deployment Impact Score0.45%0.00%100% zero-downtime
Deployment duration3.2 min5.1 min+60% (acceptable tradeoff)
Rollback success rate94%99.8%Near-perfect
Node drain disruptions2.3/week0/weekEliminated

Note: Deployments take slightly longer because we never reduce capacity (maxUnavailable: 0). This is the correct tradeoff — user experience over deployment speed.

Common Mistakes We Fixed

  1. Missing preStop hook: Without it, the pod receives SIGTERM and the endpoint is removed simultaneously. The LB still routes traffic for seconds after the pod starts shutting down.

  2. terminationGracePeriodSeconds too short: Default is 30s. If preStop sleeps for 15s and the app needs 20s to drain, the pod is killed at 30s. Set it to preStop + drain time + buffer.

  3. Readiness probe too aggressive: failureThreshold: 1 means a single slow response removes the pod from service. Use failureThreshold: 2-3 for production.

  4. No PDB on critical services: Without a PDB, kubectl drain can evict all pods of a service simultaneously. This happens during node upgrades and spot reclamation.

  5. Health endpoints that check dependencies: A readiness probe that fails when a downstream database is slow causes cascading pod removals. Readiness should check only local health.

Key Takeaways

  1. Zero-downtime is a system property: No single configuration achieves it. You need the preStop hook, proper readiness probes, PDBs, and application-level graceful shutdown working together.

  2. maxUnavailable: 0 is non-negotiable for critical services: Accept slower deployments in exchange for guaranteed capacity during updates.

  3. Measure deployment impact, not just uptime: Overall uptime can be 99.99% while individual deployments cause 0.5% failure rates. Track per-deployment impact explicitly.

  4. preStop sleep must exceed LB deregistration delay: AWS ALB takes 10-15 seconds to deregister a target. Your preStop hook must sleep longer than this.

  5. Test your graceful shutdown path: Send SIGTERM to your application locally and verify it drains active connections before exiting. Most applications do not handle this correctly without explicit implementation.

True zero-downtime deployments are achievable in Kubernetes, but only with deliberate configuration at every layer of the stack. The default settings are insufficient — treat deployment safety as a first-class engineering concern.

Comments

    No comments yet. Be the first to share your thoughts.