Mutual TLS Across 200 Microservices: Zero-Downtime Certificate Rotation at Scale
Implementing mTLS with automated certificate rotation across our Kubernetes service mesh, eliminating plaintext service-to-service communication

A network packet capture during a routine security assessment revealed something uncomfortable: 67% of our internal service-to-service traffic was unencrypted. Services within the same VPC communicated over plaintext HTTP, relying on network-level isolation for security. After the zero-trust migration exposed how permeable that network boundary actually was, encrypting all internal communication became non-negotiable. Here is how we implemented mutual TLS across 200 microservices with automated certificate rotation and zero downtime.
The Problem: Plaintext Internal Traffic
The assumption that "internal traffic is safe" collapses under scrutiny. VPC traffic can be intercepted by compromised workloads, overly permissive IAM roles can expose network interfaces, and compliance frameworks like PCI DSS require encryption regardless of network position.
Audit findings:
- 134 of 200 services communicated over plaintext HTTP internally
- Certificate management for the 66 services using TLS was manual and error-prone
- 3 production outages in 12 months caused by expired certificates
- No mutual authentication: servers had certificates, clients did not verify them
- Average certificate lifetime: 398 days (well beyond best practice of 90 days)
Architecture Decision: Istio Service Mesh
We evaluated three approaches: application-level TLS (each service manages its own certs), sidecar proxy with custom cert management, and a full service mesh. We chose Istio for its transparent mTLS enforcement, integrated certificate authority, and the operational benefits beyond just encryption.
The key insight: mTLS at the sidecar layer means applications do not need code changes. The Envoy proxy handles all TLS termination and origination transparently.
Istio Installation and Configuration
We deployed Istio with strict mTLS enforcement using a progressive rollout strategy:
apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
metadata:
name: production-mesh
spec:
profile: default
meshConfig:
# Start permissive, migrate to strict
defaultConfig:
holdApplicationUntilProxyStarts: true
accessLogFile: /dev/stdout
accessLogFormat: |
[%START_TIME%] "%REQ(:METHOD)% %REQ(X-ENVOY-ORIGINAL-PATH?:PATH)% %PROTOCOL%"
%RESPONSE_CODE% %RESPONSE_FLAGS% %BYTES_RECEIVED% %BYTES_SENT%
%DURATION% %RESP(X-ENVOY-UPSTREAM-SERVICE-TIME)%
"%REQ(X-REQUEST-ID)%" "%REQ(:AUTHORITY)%"
enableAutoMtls: true
components:
pilot:
k8s:
resources:
requests:
cpu: 500m
memory: 2Gi
limits:
cpu: 2000m
memory: 4Gi
hpaSpec:
minReplicas: 2
maxReplicas: 5
ingressGateways:
- name: istio-ingressgateway
enabled: true
k8s:
resources:
requests:
cpu: 1000m
memory: 1Gi
The enableAutoMtls: true setting is critical. It allows Istio to automatically upgrade connections to mTLS when both sides have sidecars, while gracefully falling back to plaintext for services not yet in the mesh.
Progressive mTLS Enforcement
We rolled out mTLS in three phases to avoid disruption:
Phase 1: Permissive mode (observe)
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: istio-system
spec:
mtls:
mode: PERMISSIVE # Accept both plaintext and mTLS
Phase 2: Strict per-namespace (enforce selectively)
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: strict-mtls
namespace: payments # High-security namespace first
spec:
mtls:
mode: STRICT # Reject plaintext connections
Phase 3: Mesh-wide strict enforcement
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: istio-system
spec:
mtls:
mode: STRICT # All namespaces must use mTLS
Between phases, we monitored Istio telemetry for plaintext connection attempts. Any service still sending plaintext after Phase 2 got a sidecar injection fix before Phase 3.
Certificate Rotation: The Hard Part
Istio's built-in CA (istiod) issues workload certificates with a 24-hour default lifetime. This is excellent for security but requires bulletproof rotation. We extended the default configuration with monitoring and failsafes:
apiVersion: v1
kind: ConfigMap
metadata:
name: istio-mesh-config
namespace: istio-system
data:
mesh: |
defaultConfig:
proxyMetadata:
# Certificate lifetime: 12 hours (rotate at 50% = every 6 hours)
SECRET_TTL: "12h"
SECRET_ROTATION_GRACE_PERIOD_RATIO: "0.5"
# Retry configuration for CA unavailability
SECRET_RETRY_DELAY: "5s"
SECRET_RETRY_MAX_DELAY: "60s"
We also built a certificate health monitoring system that alerts before rotation failures cascade:
import subprocess
import json
from datetime import datetime, timedelta
def check_certificate_health(namespace: str) -> dict:
"""Monitor certificate freshness across all mesh workloads.
Alerts if any certificate is within 2 hours of expiry,
indicating a rotation failure.
"""
result = subprocess.run(
["istioctl", "proxy-config", "secret", "-n", namespace, "-o", "json"],
capture_output=True, text=True
)
secrets = json.loads(result.stdout)
alerts = []
for workload in secrets:
for cert in workload.get('dynamicActiveSecrets', []):
not_after = parse_cert_expiry(cert)
time_remaining = not_after - datetime.utcnow()
if time_remaining < timedelta(hours=2):
alerts.append({
'workload': workload['name'],
'namespace': namespace,
'expires_in_minutes': int(time_remaining.total_seconds() / 60),
'severity': 'critical' if time_remaining < timedelta(hours=1) else 'warning',
})
return {
'namespace': namespace,
'total_workloads': len(secrets),
'healthy': len(secrets) - len(alerts),
'alerts': alerts,
}
Authorization Policies: mTLS + Identity
mTLS alone provides encryption and mutual authentication. Authorization policies add fine-grained access control on top:
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: payment-service-access
namespace: payments
spec:
selector:
matchLabels:
app: payment-service
action: ALLOW
rules:
- from:
- source:
# Only these services can call the payment service
principals:
- "cluster.local/ns/orders/sa/order-service"
- "cluster.local/ns/checkout/sa/checkout-service"
- "cluster.local/ns/refunds/sa/refund-service"
to:
- operation:
methods: ["POST", "GET"]
paths: ["/api/v1/payments/*", "/api/v1/refunds/*"]
- from:
- source:
principals:
- "cluster.local/ns/monitoring/sa/prometheus"
to:
- operation:
methods: ["GET"]
paths: ["/metrics", "/health"]
This policy ensures that even with network access, only authorized services can reach payment endpoints. A compromised service in the notifications namespace cannot call the payment API because its SPIFFE identity is not in the allow list.
Handling Non-Mesh Services
Not all services could join the mesh immediately. Legacy workloads, third-party integrations, and some stateful services required exceptions. We handled these with DestinationRules:
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: legacy-billing-system
namespace: billing
spec:
host: legacy-billing.billing.svc.cluster.local
trafficPolicy:
tls:
mode: DISABLE # This service cannot handle mTLS yet
---
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: legacy-billing-exception
namespace: billing
spec:
selector:
matchLabels:
app: legacy-billing
mtls:
mode: PERMISSIVE # Accept plaintext from this specific workload
# Auto-expire this exception
# Tracked via: JIRA-4521 - Migrate legacy billing to mesh
Every exception has a tracking ticket and an explicit expiration date. We review exceptions monthly and push for resolution.
Performance Impact
mTLS adds cryptographic overhead. We measured the impact carefully:
| Metric | Before mTLS | After mTLS | Delta |
|---|---|---|---|
| P50 latency (service-to-service) | 2.1ms | 2.3ms | +0.2ms |
| P99 latency (service-to-service) | 12.4ms | 13.1ms | +0.7ms |
| CPU overhead per pod (sidecar) | 0 | 15m cores | +15m |
| Memory overhead per pod (sidecar) | 0 | 45Mi | +45Mi |
| TLS handshake time (first request) | N/A | 1.8ms | One-time |
| Connection reuse rate | 72% | 94% | +22% |
The latency impact was negligible. The connection reuse improvement (from Envoy's connection pooling) actually improved overall throughput for services making many short-lived connections.
Observability: mTLS Telemetry
The mesh provides rich telemetry about TLS health:
Key metrics we monitor:
- Certificate rotation success rate (target: 100%)
- mTLS handshake failures per service
- Plaintext connection attempts (should be zero in STRICT mode)
- Certificate expiry distribution across the mesh
- Authorization policy denials by source and destination
Results After Full Rollout
| Metric | Before | After | Change |
|---|---|---|---|
| Services with encrypted internal traffic | 33% | 100% | +67% |
| Certificate-related outages (annual) | 3 | 0 | -100% |
| Certificate lifetime | 398 days | 12 hours | -99.2% |
| Unauthorized service access attempts blocked | 0 | 1,247/week | N/A |
| Mean time to revoke access | 4 hours | Immediate | -100% |
| PCI DSS encryption compliance | Partial | Full | Compliant |
Lessons Learned
Permissive mode first, always. Enabling strict mTLS without a permissive observation phase guarantees outages. We spent four weeks in permissive mode identifying services that would break.
Monitor certificate rotation, not just expiry. A rotation failure is silent until the certificate expires. Monitoring rotation success rates catches issues 6+ hours before they become outages.
Authorization policies are the real value. mTLS provides encryption, but the identity-based authorization policies are what prevent lateral movement. Deploy both together.
Sidecar resource limits matter at scale. 200 services x 3 replicas x 45Mi memory overhead = 27 GB of additional cluster memory. Budget for sidecar resources explicitly.
Conclusion
Mutual TLS across a service mesh eliminates an entire class of security vulnerabilities: plaintext interception, unauthorized service access, and certificate management failures. The performance overhead is negligible, the security improvement is substantial, and automated 12-hour certificate rotation means a compromised certificate is useful for minutes, not months. Start with permissive mode, graduate to strict enforcement namespace by namespace, and layer authorization policies on top for defense in depth.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.