Zero-Trust Network Architecture on AWS: Implementing Verified Access at Scale

How we eliminated implicit trust across 200+ microservices using AWS Verified Access, reducing lateral movement risk by 94%

#security#zero-trust#aws#networking
Cover image for the article: Zero-Trust Network Architecture on AWS: Implementing Verified Access at Scale

The traditional castle-and-moat security model assumes everything inside the VPC is trusted. After a compromised container in our staging environment traversed laterally to access production databases, we knew implicit trust had to go. Here is how we implemented zero-trust network architecture on AWS using Verified Access, identity-aware proxies, and microsegmentation across 200+ microservices.

The Problem: Implicit Trust Is a Liability

Our infrastructure had the classic setup: public subnets, private subnets, security groups, and NACLs. A VPN got engineers in, and once inside, they could reach anything the network allowed. The incident that triggered our zero-trust migration involved a supply chain compromise in a third-party library. The attacker gained code execution in a staging container, then discovered that staging and production shared a VPC peering connection with overly permissive routing.

Before zero-trust:

  • 340+ security groups with an average of 12 rules each
  • 47 VPC peering connections, many bidirectional
  • No identity verification at the network layer
  • Mean time to detect lateral movement: 6.2 hours

Architecture Overview

Our zero-trust implementation rests on three pillars: identity-first access, microsegmentation, and continuous verification.

Zero-Trust Architecture Diagram

The architecture eliminates network-level trust entirely. Every request, whether from a human engineer or a service-to-service call, must present verifiable identity claims before reaching any resource.

AWS Verified Access Configuration

AWS Verified Access became our primary control plane for human access to internal applications. We replaced our VPN entirely.

resource "aws_verifiedaccess_instance" "main" {
  description = "Zero-trust access for internal applications"

  tags = {
    Environment = "production"
    ManagedBy   = "terraform"
  }
}

resource "aws_verifiedaccess_trust_provider" "okta" {
  policy_reference_name    = "okta"
  trust_provider_type      = "user"
  user_trust_provider_type = "oidc"

  oidc_options {
    authorization_endpoint = "https://company.okta.com/oauth2/v1/authorize"
    client_id              = var.okta_client_id
    client_secret          = var.okta_client_secret
    issuer                 = "https://company.okta.com"
    scope                  = "openid profile groups"
    token_endpoint         = "https://company.okta.com/oauth2/v1/token"
    user_info_endpoint     = "https://company.okta.com/oauth2/v1/userinfo"
  }
}

resource "aws_verifiedaccess_group" "engineering" {
  verifiedaccess_instance_id = aws_verifiedaccess_instance.main.id

  policy_document = <<-EOT
    permit(principal, action, resource)
    when {
      context.okta.groups has "engineering" &&
      context.okta.email_verified == true &&
      context.http_request.http_method != "DELETE"
    };
  EOT
}

This configuration ensures that every engineer authenticating through Verified Access must belong to the engineering group in Okta and have a verified email. We further restrict destructive HTTP methods to a separate, more restrictive policy group.

Service-to-Service Zero Trust with SPIFFE/SPIRE

Human access was only half the battle. Service-to-service communication required its own identity framework. We deployed SPIRE (the SPIFFE Runtime Environment) to issue short-lived SVIDs (SPIFFE Verifiable Identity Documents) to every workload.

apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
  name: order-service
spec:
  spiffeIDTemplate: "spiffe://company.prod/ns/{{ .PodMeta.Namespace }}/sa/{{ .PodSpec.ServiceAccountName }}"
  podSelector:
    matchLabels:
      app: order-service
  ttl: "1h"
  dnsNameTemplates:
    - "{{ .PodMeta.Name }}.{{ .PodMeta.Namespace }}.svc.cluster.local"

Each service receives an identity that is cryptographically verifiable. The SVID TTL of one hour means compromised credentials have a limited blast radius. Certificate rotation happens automatically without service restarts.

Microsegmentation with Security Group Rules Engine

We built a rules engine that generates security groups from a declarative service dependency graph:

import boto3
from dataclasses import dataclass

@dataclass
class ServiceDependency:
    source: str
    destination: str
    port: int
    protocol: str = "tcp"
    justification: str = ""

def generate_security_group_rules(
    dependencies: list[ServiceDependency],
    service_registry: dict[str, str]
) -> dict:
    """Generate minimal security group rules from dependency graph."""
    rules_by_service = {}

    for dep in dependencies:
        dest_sg = service_registry[dep.destination]

        if dep.destination not in rules_by_service:
            rules_by_service[dep.destination] = []

        rules_by_service[dep.destination].append({
            "type": "ingress",
            "from_port": dep.port,
            "to_port": dep.port,
            "protocol": dep.protocol,
            "source_security_group_id": service_registry[dep.source],
            "description": f"{dep.source} -> {dep.destination}: {dep.justification}"
        })

    return rules_by_service

This approach reduced our security group rules from 4,080 to 612 while actually increasing security posture. Every rule now has a documented justification traceable to a service dependency.

Continuous Verification and Anomaly Detection

Zero-trust is not a one-time configuration. We implemented continuous verification using VPC Flow Logs, CloudTrail, and a custom anomaly detection pipeline:

Continuous Verification Pipeline

The pipeline processes approximately 2.3 billion flow log records daily, flagging connections that do not match the declared dependency graph. In the first month, it identified 23 undocumented service dependencies that needed to be either formalized or eliminated.

Results and Benchmarks

After six months of progressive rollout:

MetricBeforeAfterChange
Lateral movement risk score8.4/100.5/10-94%
Mean time to detect unauthorized access6.2 hours4.2 minutes-98.9%
Security group rules4,080612-85%
VPN-related support tickets142/month0/month-100%
Access provisioning time2.3 days14 minutes-99.6%
P99 latency overhead-+2.1msAcceptable

The latency overhead of 2.1ms at P99 comes from the identity verification step on each request. For our use case (internal services averaging 45ms response times), this was well within acceptable bounds.

Lessons Learned

Start with observability, not enforcement. We ran in audit mode for eight weeks before enabling enforcement. This surfaced 23 undocumented dependencies that would have caused outages.

Automate exception workflows. Engineers will route around security if the escape hatch is faster than the legitimate path. Our self-service exception portal with auto-expiring rules maintained velocity while preserving security.

Certificate rotation needs testing. We discovered three services that cached TLS connections indefinitely, breaking during SVID rotation. Chaos engineering (injecting early rotation) caught these before production enforcement.

VPN elimination is a morale boost. Engineers universally preferred the browser-based Verified Access flow over the old VPN client. Faster access, fewer connectivity issues, and no split-tunnel DNS headaches.

Conclusion

Zero-trust on AWS is achievable without replacing your entire infrastructure. AWS Verified Access handles human access, SPIFFE/SPIRE handles workload identity, and a declarative dependency graph drives microsegmentation. The 94% reduction in lateral movement risk validated the investment. Start with audit mode, map your dependencies, and migrate incrementally. The castle-and-moat model served its era, but modern threat landscapes demand identity verification at every boundary.

Comments

    No comments yet. Be the first to share your thoughts.