Mutual TLS Across 200 Microservices: Zero-Downtime Certificate Rotation at Scale

Implementing mTLS with automated certificate rotation across our Kubernetes service mesh, eliminating plaintext service-to-service communication

#service-mesh#mtls#security#kubernetes
Cover image for the article: Mutual TLS Across 200 Microservices: Zero-Downtime Certificate Rotation at Scale

A network packet capture during a routine security assessment revealed something uncomfortable: 67% of our internal service-to-service traffic was unencrypted. Services within the same VPC communicated over plaintext HTTP, relying on network-level isolation for security. After the zero-trust migration exposed how permeable that network boundary actually was, encrypting all internal communication became non-negotiable. Here is how we implemented mutual TLS across 200 microservices with automated certificate rotation and zero downtime.

The Problem: Plaintext Internal Traffic

The assumption that "internal traffic is safe" collapses under scrutiny. VPC traffic can be intercepted by compromised workloads, overly permissive IAM roles can expose network interfaces, and compliance frameworks like PCI DSS require encryption regardless of network position.

Audit findings:

  • 134 of 200 services communicated over plaintext HTTP internally
  • Certificate management for the 66 services using TLS was manual and error-prone
  • 3 production outages in 12 months caused by expired certificates
  • No mutual authentication: servers had certificates, clients did not verify them
  • Average certificate lifetime: 398 days (well beyond best practice of 90 days)

Architecture Decision: Istio Service Mesh

We evaluated three approaches: application-level TLS (each service manages its own certs), sidecar proxy with custom cert management, and a full service mesh. We chose Istio for its transparent mTLS enforcement, integrated certificate authority, and the operational benefits beyond just encryption.

Service Mesh mTLS Architecture

The key insight: mTLS at the sidecar layer means applications do not need code changes. The Envoy proxy handles all TLS termination and origination transparently.

Istio Installation and Configuration

We deployed Istio with strict mTLS enforcement using a progressive rollout strategy:

apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
metadata:
  name: production-mesh
spec:
  profile: default
  meshConfig:
    # Start permissive, migrate to strict
    defaultConfig:
      holdApplicationUntilProxyStarts: true
    accessLogFile: /dev/stdout
    accessLogFormat: |
      [%START_TIME%] "%REQ(:METHOD)% %REQ(X-ENVOY-ORIGINAL-PATH?:PATH)% %PROTOCOL%"
      %RESPONSE_CODE% %RESPONSE_FLAGS% %BYTES_RECEIVED% %BYTES_SENT%
      %DURATION% %RESP(X-ENVOY-UPSTREAM-SERVICE-TIME)%
      "%REQ(X-REQUEST-ID)%" "%REQ(:AUTHORITY)%"
    enableAutoMtls: true
  components:
    pilot:
      k8s:
        resources:
          requests:
            cpu: 500m
            memory: 2Gi
          limits:
            cpu: 2000m
            memory: 4Gi
        hpaSpec:
          minReplicas: 2
          maxReplicas: 5
    ingressGateways:
      - name: istio-ingressgateway
        enabled: true
        k8s:
          resources:
            requests:
              cpu: 1000m
              memory: 1Gi

The enableAutoMtls: true setting is critical. It allows Istio to automatically upgrade connections to mTLS when both sides have sidecars, while gracefully falling back to plaintext for services not yet in the mesh.

Progressive mTLS Enforcement

We rolled out mTLS in three phases to avoid disruption:

Phase 1: Permissive mode (observe)

apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: istio-system
spec:
  mtls:
    mode: PERMISSIVE  # Accept both plaintext and mTLS

Phase 2: Strict per-namespace (enforce selectively)

apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: strict-mtls
  namespace: payments  # High-security namespace first
spec:
  mtls:
    mode: STRICT  # Reject plaintext connections

Phase 3: Mesh-wide strict enforcement

apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: istio-system
spec:
  mtls:
    mode: STRICT  # All namespaces must use mTLS

Between phases, we monitored Istio telemetry for plaintext connection attempts. Any service still sending plaintext after Phase 2 got a sidecar injection fix before Phase 3.

Certificate Rotation: The Hard Part

Istio's built-in CA (istiod) issues workload certificates with a 24-hour default lifetime. This is excellent for security but requires bulletproof rotation. We extended the default configuration with monitoring and failsafes:

apiVersion: v1
kind: ConfigMap
metadata:
  name: istio-mesh-config
  namespace: istio-system
data:
  mesh: |
    defaultConfig:
      proxyMetadata:
        # Certificate lifetime: 12 hours (rotate at 50% = every 6 hours)
        SECRET_TTL: "12h"
        SECRET_ROTATION_GRACE_PERIOD_RATIO: "0.5"
        # Retry configuration for CA unavailability
        SECRET_RETRY_DELAY: "5s"
        SECRET_RETRY_MAX_DELAY: "60s"

We also built a certificate health monitoring system that alerts before rotation failures cascade:

import subprocess
import json
from datetime import datetime, timedelta

def check_certificate_health(namespace: str) -> dict:
    """Monitor certificate freshness across all mesh workloads.
    
    Alerts if any certificate is within 2 hours of expiry,
    indicating a rotation failure.
    """
    result = subprocess.run(
        ["istioctl", "proxy-config", "secret", "-n", namespace, "-o", "json"],
        capture_output=True, text=True
    )
    secrets = json.loads(result.stdout)
    
    alerts = []
    for workload in secrets:
        for cert in workload.get('dynamicActiveSecrets', []):
            not_after = parse_cert_expiry(cert)
            time_remaining = not_after - datetime.utcnow()
            
            if time_remaining < timedelta(hours=2):
                alerts.append({
                    'workload': workload['name'],
                    'namespace': namespace,
                    'expires_in_minutes': int(time_remaining.total_seconds() / 60),
                    'severity': 'critical' if time_remaining < timedelta(hours=1) else 'warning',
                })
    
    return {
        'namespace': namespace,
        'total_workloads': len(secrets),
        'healthy': len(secrets) - len(alerts),
        'alerts': alerts,
    }

Authorization Policies: mTLS + Identity

mTLS alone provides encryption and mutual authentication. Authorization policies add fine-grained access control on top:

apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
  name: payment-service-access
  namespace: payments
spec:
  selector:
    matchLabels:
      app: payment-service
  action: ALLOW
  rules:
    - from:
        - source:
            # Only these services can call the payment service
            principals:
              - "cluster.local/ns/orders/sa/order-service"
              - "cluster.local/ns/checkout/sa/checkout-service"
              - "cluster.local/ns/refunds/sa/refund-service"
      to:
        - operation:
            methods: ["POST", "GET"]
            paths: ["/api/v1/payments/*", "/api/v1/refunds/*"]
    - from:
        - source:
            principals:
              - "cluster.local/ns/monitoring/sa/prometheus"
      to:
        - operation:
            methods: ["GET"]
            paths: ["/metrics", "/health"]

This policy ensures that even with network access, only authorized services can reach payment endpoints. A compromised service in the notifications namespace cannot call the payment API because its SPIFFE identity is not in the allow list.

Handling Non-Mesh Services

Not all services could join the mesh immediately. Legacy workloads, third-party integrations, and some stateful services required exceptions. We handled these with DestinationRules:

apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: legacy-billing-system
  namespace: billing
spec:
  host: legacy-billing.billing.svc.cluster.local
  trafficPolicy:
    tls:
      mode: DISABLE  # This service cannot handle mTLS yet
---
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: legacy-billing-exception
  namespace: billing
spec:
  selector:
    matchLabels:
      app: legacy-billing
  mtls:
    mode: PERMISSIVE  # Accept plaintext from this specific workload
  # Auto-expire this exception
  # Tracked via: JIRA-4521 - Migrate legacy billing to mesh

Every exception has a tracking ticket and an explicit expiration date. We review exceptions monthly and push for resolution.

Performance Impact

mTLS adds cryptographic overhead. We measured the impact carefully:

MetricBefore mTLSAfter mTLSDelta
P50 latency (service-to-service)2.1ms2.3ms+0.2ms
P99 latency (service-to-service)12.4ms13.1ms+0.7ms
CPU overhead per pod (sidecar)015m cores+15m
Memory overhead per pod (sidecar)045Mi+45Mi
TLS handshake time (first request)N/A1.8msOne-time
Connection reuse rate72%94%+22%

The latency impact was negligible. The connection reuse improvement (from Envoy's connection pooling) actually improved overall throughput for services making many short-lived connections.

Observability: mTLS Telemetry

The mesh provides rich telemetry about TLS health:

mTLS Observability Dashboard

Key metrics we monitor:

  • Certificate rotation success rate (target: 100%)
  • mTLS handshake failures per service
  • Plaintext connection attempts (should be zero in STRICT mode)
  • Certificate expiry distribution across the mesh
  • Authorization policy denials by source and destination

Results After Full Rollout

MetricBeforeAfterChange
Services with encrypted internal traffic33%100%+67%
Certificate-related outages (annual)30-100%
Certificate lifetime398 days12 hours-99.2%
Unauthorized service access attempts blocked01,247/weekN/A
Mean time to revoke access4 hoursImmediate-100%
PCI DSS encryption compliancePartialFullCompliant

Lessons Learned

Permissive mode first, always. Enabling strict mTLS without a permissive observation phase guarantees outages. We spent four weeks in permissive mode identifying services that would break.

Monitor certificate rotation, not just expiry. A rotation failure is silent until the certificate expires. Monitoring rotation success rates catches issues 6+ hours before they become outages.

Authorization policies are the real value. mTLS provides encryption, but the identity-based authorization policies are what prevent lateral movement. Deploy both together.

Sidecar resource limits matter at scale. 200 services x 3 replicas x 45Mi memory overhead = 27 GB of additional cluster memory. Budget for sidecar resources explicitly.

Conclusion

Mutual TLS across a service mesh eliminates an entire class of security vulnerabilities: plaintext interception, unauthorized service access, and certificate management failures. The performance overhead is negligible, the security improvement is substantial, and automated 12-hour certificate rotation means a compromised certificate is useful for minutes, not months. Start with permissive mode, graduate to strict enforcement namespace by namespace, and layer authorization policies on top for defense in depth.

Comments

    No comments yet. Be the first to share your thoughts.