Zero-Trust Network Architecture on AWS: Implementing Verified Access at Scale
How we eliminated implicit trust across 200+ microservices using AWS Verified Access, reducing lateral movement risk by 94%

The traditional castle-and-moat security model assumes everything inside the VPC is trusted. After a compromised container in our staging environment traversed laterally to access production databases, we knew implicit trust had to go. Here is how we implemented zero-trust network architecture on AWS using Verified Access, identity-aware proxies, and microsegmentation across 200+ microservices.
The Problem: Implicit Trust Is a Liability
Our infrastructure had the classic setup: public subnets, private subnets, security groups, and NACLs. A VPN got engineers in, and once inside, they could reach anything the network allowed. The incident that triggered our zero-trust migration involved a supply chain compromise in a third-party library. The attacker gained code execution in a staging container, then discovered that staging and production shared a VPC peering connection with overly permissive routing.
Before zero-trust:
- 340+ security groups with an average of 12 rules each
- 47 VPC peering connections, many bidirectional
- No identity verification at the network layer
- Mean time to detect lateral movement: 6.2 hours
Architecture Overview
Our zero-trust implementation rests on three pillars: identity-first access, microsegmentation, and continuous verification.
The architecture eliminates network-level trust entirely. Every request, whether from a human engineer or a service-to-service call, must present verifiable identity claims before reaching any resource.
AWS Verified Access Configuration
AWS Verified Access became our primary control plane for human access to internal applications. We replaced our VPN entirely.
resource "aws_verifiedaccess_instance" "main" {
description = "Zero-trust access for internal applications"
tags = {
Environment = "production"
ManagedBy = "terraform"
}
}
resource "aws_verifiedaccess_trust_provider" "okta" {
policy_reference_name = "okta"
trust_provider_type = "user"
user_trust_provider_type = "oidc"
oidc_options {
authorization_endpoint = "https://company.okta.com/oauth2/v1/authorize"
client_id = var.okta_client_id
client_secret = var.okta_client_secret
issuer = "https://company.okta.com"
scope = "openid profile groups"
token_endpoint = "https://company.okta.com/oauth2/v1/token"
user_info_endpoint = "https://company.okta.com/oauth2/v1/userinfo"
}
}
resource "aws_verifiedaccess_group" "engineering" {
verifiedaccess_instance_id = aws_verifiedaccess_instance.main.id
policy_document = <<-EOT
permit(principal, action, resource)
when {
context.okta.groups has "engineering" &&
context.okta.email_verified == true &&
context.http_request.http_method != "DELETE"
};
EOT
}
This configuration ensures that every engineer authenticating through Verified Access must belong to the engineering group in Okta and have a verified email. We further restrict destructive HTTP methods to a separate, more restrictive policy group.
Service-to-Service Zero Trust with SPIFFE/SPIRE
Human access was only half the battle. Service-to-service communication required its own identity framework. We deployed SPIRE (the SPIFFE Runtime Environment) to issue short-lived SVIDs (SPIFFE Verifiable Identity Documents) to every workload.
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
name: order-service
spec:
spiffeIDTemplate: "spiffe://company.prod/ns/{{ .PodMeta.Namespace }}/sa/{{ .PodSpec.ServiceAccountName }}"
podSelector:
matchLabels:
app: order-service
ttl: "1h"
dnsNameTemplates:
- "{{ .PodMeta.Name }}.{{ .PodMeta.Namespace }}.svc.cluster.local"
Each service receives an identity that is cryptographically verifiable. The SVID TTL of one hour means compromised credentials have a limited blast radius. Certificate rotation happens automatically without service restarts.
Microsegmentation with Security Group Rules Engine
We built a rules engine that generates security groups from a declarative service dependency graph:
import boto3
from dataclasses import dataclass
@dataclass
class ServiceDependency:
source: str
destination: str
port: int
protocol: str = "tcp"
justification: str = ""
def generate_security_group_rules(
dependencies: list[ServiceDependency],
service_registry: dict[str, str]
) -> dict:
"""Generate minimal security group rules from dependency graph."""
rules_by_service = {}
for dep in dependencies:
dest_sg = service_registry[dep.destination]
if dep.destination not in rules_by_service:
rules_by_service[dep.destination] = []
rules_by_service[dep.destination].append({
"type": "ingress",
"from_port": dep.port,
"to_port": dep.port,
"protocol": dep.protocol,
"source_security_group_id": service_registry[dep.source],
"description": f"{dep.source} -> {dep.destination}: {dep.justification}"
})
return rules_by_service
This approach reduced our security group rules from 4,080 to 612 while actually increasing security posture. Every rule now has a documented justification traceable to a service dependency.
Continuous Verification and Anomaly Detection
Zero-trust is not a one-time configuration. We implemented continuous verification using VPC Flow Logs, CloudTrail, and a custom anomaly detection pipeline:
The pipeline processes approximately 2.3 billion flow log records daily, flagging connections that do not match the declared dependency graph. In the first month, it identified 23 undocumented service dependencies that needed to be either formalized or eliminated.
Results and Benchmarks
After six months of progressive rollout:
| Metric | Before | After | Change |
|---|---|---|---|
| Lateral movement risk score | 8.4/10 | 0.5/10 | -94% |
| Mean time to detect unauthorized access | 6.2 hours | 4.2 minutes | -98.9% |
| Security group rules | 4,080 | 612 | -85% |
| VPN-related support tickets | 142/month | 0/month | -100% |
| Access provisioning time | 2.3 days | 14 minutes | -99.6% |
| P99 latency overhead | - | +2.1ms | Acceptable |
The latency overhead of 2.1ms at P99 comes from the identity verification step on each request. For our use case (internal services averaging 45ms response times), this was well within acceptable bounds.
Lessons Learned
Start with observability, not enforcement. We ran in audit mode for eight weeks before enabling enforcement. This surfaced 23 undocumented dependencies that would have caused outages.
Automate exception workflows. Engineers will route around security if the escape hatch is faster than the legitimate path. Our self-service exception portal with auto-expiring rules maintained velocity while preserving security.
Certificate rotation needs testing. We discovered three services that cached TLS connections indefinitely, breaking during SVID rotation. Chaos engineering (injecting early rotation) caught these before production enforcement.
VPN elimination is a morale boost. Engineers universally preferred the browser-based Verified Access flow over the old VPN client. Faster access, fewer connectivity issues, and no split-tunnel DNS headaches.
Conclusion
Zero-trust on AWS is achievable without replacing your entire infrastructure. AWS Verified Access handles human access, SPIFFE/SPIRE handles workload identity, and a declarative dependency graph drives microsegmentation. The 94% reduction in lateral movement risk validated the investment. Start with audit mode, map your dependencies, and migrate incrementally. The castle-and-moat model served its era, but modern threat landscapes demand identity verification at every boundary.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.