Kubernetes Network Policies for Zero-Trust: Blocking Lateral Movement in 200-Pod Clusters
How we implemented zero-trust networking in a 200-pod Kubernetes cluster using network policies, reducing our attack surface by 94% without breaking services.

In a default Kubernetes cluster, every pod can talk to every other pod. That's terrifying. When we ran a penetration test on our 200-pod production cluster, the assessment team pivoted from a compromised logging sidecar to our payment service in three hops. No firewall, no segmentation, no resistance. This is how we implemented zero-trust networking that reduced our reachable attack surface by 94%.
The Default Kubernetes Network Model
Out of the box, Kubernetes networking is flat and permissive:
- Every pod can reach every other pod across all namespaces
- Every pod can reach external endpoints
- There are no ingress restrictions
- Service discovery is global — any pod can resolve any service DNS
This is great for development velocity. It's catastrophic for security. In a 200-pod cluster with 35 services, the default network model creates 200 × 199 = 39,800 possible pod-to-pod communication paths. Our services actually needed fewer than 400 of those paths.
| Metric | Before Network Policies | After Network Policies |
|---|---|---|
| Possible communication paths | 39,800 | 2,340 (allowed) |
| Actual required paths | ~400 | ~400 |
| Attack surface (reachable pods from any pod) | 199 (100%) | 12 avg (6%) |
| Lateral movement hops to critical services | 1-2 | Blocked |
Choosing a CNI: Calico vs Cilium
Network policies require a CNI that enforces them. We evaluated the two leading options:
| Feature | Calico | Cilium |
|---|---|---|
| Network policy enforcement | iptables-based | eBPF-based |
| Performance at scale (200+ pods) | Good | Excellent |
| L7 policy support | GlobalNetworkPolicy | HTTP-aware policies |
| DNS-based policies | Yes (Calico Enterprise) | Yes (native) |
| Observability | Flow logs | Hubble (built-in) |
| Encryption (WireGuard) | Yes | Yes |
| Resource overhead | Medium | Lower at scale |
We chose Cilium for its eBPF-based enforcement (lower performance overhead at our scale), native L7 policy support, and Hubble observability. The ability to write policies based on HTTP methods and paths — not just L3/L4 — was decisive for our API-driven architecture.
The Implementation Strategy: Default-Deny First
The cardinal rule of zero-trust networking: start with deny-all, then explicitly allow required traffic. We rolled this out namespace by namespace over four weeks.
Step 1: Default Deny All Traffic
# default-deny.yaml - Apply to every namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {} # Matches all pods in namespace
policyTypes:
- Ingress
- Egress
# No ingress/egress rules = deny everything
Step 2: Allow DNS (Critical — Everything Breaks Without This)
# allow-dns.yaml - Must apply before default-deny or services can't resolve
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns
namespace: production
spec:
podSelector: {}
policyTypes:
- Egress
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
Step 3: Service-Specific Allow Rules
Here's where the real work happens. Each service gets explicit policies for its dependencies:
# payment-service-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: payment-service-policy
namespace: production
spec:
podSelector:
matchLabels:
app: payment-service
policyTypes:
- Ingress
- Egress
ingress:
# Only the API gateway can reach payment service
- from:
- podSelector:
matchLabels:
app: api-gateway
ports:
- protocol: TCP
port: 8080
# Health checks from the ingress controller
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: ingress-nginx
ports:
- protocol: TCP
port: 8080
egress:
# Payment service can reach the database
- to:
- podSelector:
matchLabels:
app: postgres
tier: payments-db
ports:
- protocol: TCP
port: 5432
# Payment service can reach Stripe (external)
- to:
- ipBlock:
cidr: 0.0.0.0/0
ports:
- protocol: TCP
port: 443
Cilium L7 Policies: HTTP-Aware Segmentation
Standard Kubernetes NetworkPolicies only work at L3/L4 (IP + port). Cilium extends this to L7, allowing us to restrict which HTTP methods and paths a service can access:
# Cilium L7 policy - restrict API gateway to specific endpoints
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: api-gateway-l7-policy
namespace: production
spec:
endpointSelector:
matchLabels:
app: user-service
ingress:
- fromEndpoints:
- matchLabels:
app: api-gateway
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
# API gateway can only call these specific endpoints
- method: GET
path: "/api/v1/users/.*"
- method: POST
path: "/api/v1/users"
- method: GET
path: "/health"
# Explicitly NOT allowing DELETE or admin endpoints
This means even if the API gateway is compromised, the attacker cannot call administrative endpoints or destructive operations on downstream services.
Observability: Knowing What's Happening
You cannot secure what you cannot see. Cilium Hubble provides real-time flow visibility:
# Monitor all dropped traffic (policy violations)
hubble observe --verdict DROPPED --namespace production
# Watch traffic to payment service specifically
hubble observe --to-label app=payment-service --namespace production
# Export flows for analysis
hubble observe --output json | jq 'select(.verdict == "DROPPED")'
We built a dashboard showing:
- All denied flows (potential misconfigurations OR attack attempts)
- Traffic graph of actual service-to-service communication
- Policy coverage percentage per namespace
# Prometheus metrics from Cilium for alerting
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: network-policy-alerts
spec:
groups:
- name: network-policy-violations
rules:
- alert: HighPolicyDeniedRate
expr: |
rate(cilium_drop_count_total{reason="POLICY_DENIED"}[5m]) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "High rate of network policy denials"
description: "{{ $labels.namespace }}/{{ $labels.pod }} is being denied at {{ $value }} requests/sec"
The Rollout Process (Without Breaking Production)
Rolling out network policies to a running cluster is nerve-wracking. Here's our four-week process:
Week 1: Audit mode — Deploy Cilium policies in "audit" mode (log but don't enforce). Collect all actual traffic flows.
Week 2: Generate policies — Use observed flows to auto-generate baseline policies with cilium connectivity test and Hubble flow logs.
Week 3: Enforce non-critical namespaces — Apply policies to staging and internal tools first. Fix any broken connections.
Week 4: Enforce production — Apply to production with a rollback plan (remove default-deny) ready.
// Policy generator from Hubble flow logs
interface FlowLog {
source: { labels: Record<string, string> };
destination: { labels: Record<string, string>; port: number };
verdict: 'FORWARDED' | 'DROPPED';
}
function generatePoliciesFromFlows(flows: FlowLog[]): NetworkPolicy[] {
const serviceFlows = new Map<string, Set<string>>();
for (const flow of flows) {
if (flow.verdict !== 'FORWARDED') continue;
const source = flow.source.labels['app'];
const dest = `${flow.destination.labels['app']}:${flow.destination.port}`;
if (!serviceFlows.has(source)) serviceFlows.set(source, new Set());
serviceFlows.get(source)!.add(dest);
}
// Generate minimal allow policies from observed traffic
return Array.from(serviceFlows.entries()).map(([source, destinations]) =>
buildNetworkPolicy(source, destinations)
);
}
Results: Penetration Test Comparison
We ran identical penetration tests before and after:
| Attack Scenario | Before | After |
|---|---|---|
| Lateral movement from compromised pod | 3 hops to payment service | Blocked at first hop |
| Data exfiltration via DNS tunneling | Unrestricted | Blocked (DNS only to kube-dns) |
| Port scanning from within cluster | All 39,800 paths reachable | 6% reachable (12 pods avg) |
| Unauthorized API calls | Any pod → any endpoint | L7 method/path restricted |
| Egress to C2 server | Unrestricted | Blocked (allowlist only) |
Performance Impact
A common concern with network policies is performance overhead. Our measurements:
| Metric | Without Policies | With Cilium Policies | Impact |
|---|---|---|---|
| P50 inter-service latency | 0.8ms | 0.85ms | +6% |
| P99 inter-service latency | 4.2ms | 4.5ms | +7% |
| CPU overhead (per node) | Baseline | +2.3% | Minimal |
| Memory overhead (per node) | Baseline | +180MB | Acceptable |
| Max throughput | 42K req/s | 41.2K req/s | -2% |
The eBPF-based enforcement adds negligible overhead. IPtables-based solutions (Calico in default mode) would show higher impact at this scale.
Key Takeaways
-
Default-deny is non-negotiable — if you're running Kubernetes in production without network policies, every pod is one exploit away from being a pivot point.
-
Start with observability, not enforcement — run Hubble or Calico flow logs for two weeks before writing any policies. Let real traffic patterns inform your rules.
-
L7 policies provide defense-in-depth — restricting not just which services can communicate but what operations they can perform dramatically reduces blast radius.
-
DNS policies are the first thing to get right — blocking DNS breaks everything. Allow kube-dns before applying default-deny.
-
eBPF-based enforcement scales better — at 200+ pods, the difference between eBPF (Cilium) and iptables (default Calico) becomes measurable.
-
Network policies are a security control, not a silver bullet — combine with mTLS (service mesh), RBAC, and pod security standards for comprehensive zero-trust.
The flat Kubernetes network is a liability. Network policies transform it into a segmented, auditable, zero-trust environment where compromising one service doesn't mean compromising all of them. The two weeks of implementation work pays dividends every time a security audit passes without findings.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.