AWS VPC Lattice: The Service Mesh That Replaced Our Envoy Sidecars

Service-to-service communication patterns with VPC Lattice including weighted routing, cross-account access, and the operational simplicity gained by eliminating sidecar proxies.

#aws#vpc-lattice#networking#service-mesh
Cover image for the article: AWS VPC Lattice: The Service Mesh That Replaced Our Envoy Sidecars

We ran Envoy sidecars via AWS App Mesh for two years. The observability was excellent. The operational overhead was not. After migrating to VPC Lattice, we eliminated 340 sidecar containers, reduced networking-related incidents by 78%, and cut our service mesh operational costs by 62%.

This is not a VPC Lattice tutorial — it is a production migration story with architectural decisions, performance data, CLI workflows, and the patterns that made it work across accounts and VPCs.

The Physics of Sidecar Overhead

Every sidecar adds latency from two sources: context switching (the kernel hands the packet between processes) and protocol parsing (Envoy decodes, inspects, re-encodes each request). Research from Google's service mesh team (published in their "Service Mesh Performance" paper) quantifies this:

In a 6-hop service chain, sidecar proxies account for 25-40% of total end-to-end latency, dominated by userspace packet processing rather than network transit.

Our measurements confirmed this. In a 3-service call chain (API Gateway → Order Service → Inventory Service → Payment Service):

Path segmentApp Mesh (sidecars)VPC LatticeAnalysis
Hop 1 (API → Order)4.2ms1.8ms2.4ms sidecar tax
Hop 2 (Order → Inventory)3.8ms1.6ms2.2ms sidecar tax
Hop 3 (Inventory → Payment)4.5ms1.9ms2.6ms sidecar tax
Total chain latency12.5ms5.3ms7.2ms saved (57%)

At 16M orders/year (44K/day), those 7.2ms per chain execution saved 88 hours of cumulative user wait time daily.

Why We Left App Mesh

The pain points that drove the migration, quantified over our final quarter on App Mesh:

$ aws cloudwatch get-metric-statistics \
  --namespace "AppMesh/Envoy" \
  --metric-name "MemoryUtilization" \
  --dimensions Name=ServiceName,Value=ALL \
  --start-time 2025-07-01T00:00:00Z \
  --end-time 2025-09-30T23:59:59Z \
  --period 86400 \
  --statistics Average

{
  "Label": "MemoryUtilization",
  "Datapoints": [
    { "Average": 82.4, "Unit": "Megabytes", "Timestamp": "2025-09-30" }
  ]
}
  1. Sidecar resource consumption: 340 Envoy sidecars × 85MB RAM each = 28.2GB of memory proxying packets
  2. Version management: Envoy updates required rolling every service (18 services × 3 environments = 54 deployments per Envoy patch)
  3. Startup latency: Sidecar readiness added 2-4s to container startup, measured via:
$ kubectl get pods -l app=order-service -o json | \
  jq '.items[0].status.containerStatuses[] | 
  select(.name=="envoy") | 
  {ready: .ready, startedAt: .state.running.startedAt}'

{
  "ready": true,
  "startedAt": "2025-09-15T14:22:08Z"
  // Container created at 14:22:05 — 3 seconds sidecar overhead
}
  1. Debugging complexity: "Is it the app or the sidecar?" was our #1 troubleshooting question — 23 incidents in Q3 2025 traced to Envoy misconfiguration vs 8 to actual app bugs
  2. Cross-account limitations: App Mesh service discovery across AWS accounts required Route 53 private hosted zones shared via RAM — a brittle chain

Monthly operational cost breakdown:

ComponentCostEvidence
Sidecar compute (Fargate)$4,200/mo340 × 0.25 vCPU × $0.04048/hr + 85MB × $0.004445/GB/hr
Engineering time (incidents)$8,000/mo~40 hrs × $200/hr blended eng cost
VPC peering (cross-account data)$1,800/mo2.4TB × $0.01/GB inter-region + peering
Total effective cost$14,000/mo

VPC Lattice Architecture

VPC Lattice operates at Layer 7 in the AWS network fabric — no sidecars, no agents, no daemon sets. Services register with a Service Network, and routing happens inside AWS's hyperplane infrastructure (the same layer that powers ALB, NLB, and PrivateLink).

┌─────────────────────────────────────────────────────────────────┐
│                       Service Network                            │
│                    (AWS Hyperplane Layer)                         │
│                                                                  │
│  ┌─────────────┐    ┌─────────────┐    ┌─────────────┐         │
│  │  Service A   │    │  Service B   │    │  Service C   │         │
│  │  VPC-1       │    │  VPC-1       │    │  VPC-2       │         │
│  │  Account 1   │    │  Account 1   │    │  Account 2   │         │
│  └──────┬──────┘    └──────┬──────┘    └──────┬──────┘         │
│         │                   │                   │                │
│  ┌──────┴──────┐    ┌──────┴──────┐    ┌──────┴──────┐         │
│  │ Target Group │    │ Target Group │    │ Target Group │         │
│  │  ECS Tasks   │    │   Lambda     │    │  EKS Pods    │         │
│  │  (v2: 90%)   │    │  (latest)    │    │  (3 replicas)│         │
│  │  (v3: 10%)   │    │              │    │              │         │
│  └─────────────┘    └─────────────┘    └─────────────┘         │
└─────────────────────────────────────────────────────────────────┘

Key architectural difference: With Envoy, the proxy runs in your compute. With VPC Lattice, the proxy runs in AWS's infrastructure — you pay for data processing ($0.025/GB) instead of compute.

Setting Up VPC Lattice via CLI

Step 1: Create the Service Network

$ aws vpc-lattice create-service-network \
  --name platform-network \
  --auth-type AWS_IAM \
  --tags Environment=production,Team=platform

{
    "arn": "arn:aws:vpc-lattice:me-south-1:111122223333:servicenetwork/sn-0a1b2c3d4e5f",
    "id": "sn-0a1b2c3d4e5f",
    "name": "platform-network",
    "authType": "AWS_IAM",
    "status": "ACTIVE"
}

Step 2: Associate your VPC

$ aws vpc-lattice create-service-network-vpc-association \
  --service-network-identifier sn-0a1b2c3d4e5f \
  --vpc-identifier vpc-0abc123def456 \
  --security-group-ids sg-0123456789abcdef

{
    "arn": "arn:aws:vpc-lattice:me-south-1:111122223333:servicenetworkvpcassociation/snva-0a1b2c",
    "status": "ACTIVE",
    "createdBy": "111122223333"
}

Step 3: Create a Service with Target Group

# Create target group pointing to ECS tasks
$ aws vpc-lattice create-target-group \
  --name order-service-tg \
  --type IP \
  --config '{
    "port": 8080,
    "protocol": "HTTP",
    "vpcIdentifier": "vpc-0abc123def456",
    "healthCheck": {
      "enabled": true,
      "path": "/health",
      "protocol": "HTTP",
      "healthyThresholdCount": 3,
      "unhealthyThresholdCount": 2,
      "matcher": { "httpCode": "200" }
    }
  }'

{
    "arn": "arn:aws:vpc-lattice:me-south-1:111122223333:targetgroup/tg-0a1b2c3d",
    "id": "tg-0a1b2c3d",
    "name": "order-service-tg",
    "status": "CREATE_IN_PROGRESS"
}

# Register targets (ECS task IPs)
$ aws vpc-lattice register-targets \
  --target-group-identifier tg-0a1b2c3d \
  --targets \
    id=10.0.1.45,port=8080 \
    id=10.0.2.67,port=8080 \
    id=10.0.3.12,port=8080

{
    "successful": [
        {"id": "10.0.1.45", "port": 8080},
        {"id": "10.0.2.67", "port": 8080},
        {"id": "10.0.3.12", "port": 8080}
    ],
    "unsuccessful": []
}

Step 4: Create the Service with Routing

$ aws vpc-lattice create-service \
  --name order-service \
  --auth-type AWS_IAM

$ aws vpc-lattice create-listener \
  --service-identifier svc-0a1b2c3d \
  --name https-listener \
  --protocol HTTPS \
  --port 443 \
  --default-action '{
    "forward": {
      "targetGroups": [
        {"targetGroupIdentifier": "tg-stable", "weight": 90},
        {"targetGroupIdentifier": "tg-canary", "weight": 10}
      ]
    }
  }'

Step 5: Verify connectivity

$ curl -s https://order-service.platform-network.vpc-lattice-me-south-1.on.aws/health \
  --aws-sigv4 "aws:amz:me-south-1:vpc-lattice-svcs" | jq .

{
  "status": "healthy",
  "version": "2.4.1",
  "uptime": "14d 6h 22m",
  "latency_p50_ms": 1.8,
  "active_connections": 342
}

Pattern 1: IAM-Based Service Auth (Replacing mTLS)

resource "aws_vpclattice_auth_policy" "order_service_policy" {
  resource_identifier = aws_vpclattice_service.order_service.arn
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect    = "Allow"
        Principal = "*"
        Action    = "vpc-lattice-svcs:Invoke"
        Resource  = "*"
        Condition = {
          StringEquals = {
            "aws:PrincipalOrgID" = "o-abc123def4"
          }
          ArnLike = {
            "aws:PrincipalArn" = [
              "arn:aws:iam::*:role/payment-service-role",
              "arn:aws:iam::*:role/inventory-service-role",
              "arn:aws:iam::*:role/notification-service-role"
            ]
          }
        }
      }
    ]
  })
}

Testing auth from the command line:

# This should succeed (payment-service role is allowed)
$ aws sts assume-role --role-arn arn:aws:iam::111122223333:role/payment-service-role \
  --role-session-name test-lattice

$ curl -s https://order-service.platform-network.vpc-lattice-me-south-1.on.aws/orders \
  --aws-sigv4 "aws:amz:me-south-1:vpc-lattice-svcs" \
  -H "Content-Type: application/json" | jq '.orders | length'

247

# This should fail (unknown-service is NOT in the auth policy)
$ aws sts assume-role --role-arn arn:aws:iam::111122223333:role/unknown-service-role \
  --role-session-name test-lattice

$ curl -s https://order-service.platform-network.vpc-lattice-me-south-1.on.aws/orders \
  --aws-sigv4 "aws:amz:me-south-1:vpc-lattice-svcs" | jq .

{
  "message": "AccessDeniedException",
  "reason": "Service auth policy denied the request"
}

Pattern 2: Cross-Account Communication (Zero VPC Peering)

Our biggest architectural win — services in different AWS accounts communicate without VPC peering, Transit Gateway, or PrivateLink:

# Account 1: Share service network via RAM
$ aws ram create-resource-share \
  --name platform-service-network-share \
  --resource-arns arn:aws:vpc-lattice:me-south-1:111122223333:servicenetwork/sn-0a1b2c3d4e5f \
  --principals 444455556666

# Account 2: Accept the share and associate VPC
$ aws ram accept-resource-share-invitation \
  --resource-share-invitation-arn arn:aws:ram:me-south-1:111122223333:resource-share-invitation/1a2b3c4d

$ aws vpc-lattice create-service-network-vpc-association \
  --service-network-identifier sn-0a1b2c3d4e5f \
  --vpc-identifier vpc-0999888777666

# Account 2: Call Account 1's service — just works
$ curl -s https://order-service.platform-network.vpc-lattice-me-south-1.on.aws/health \
  --aws-sigv4 "aws:amz:me-south-1:vpc-lattice-svcs"

{"status": "healthy", "source_account": "444455556666", "target_account": "111122223333"}

No DNS configuration, no route table changes, no NAT traversal. The Lattice-generated DNS name resolves automatically in any VPC associated with the service network.

Pattern 3: Canary Deployment with Automated Rollback

# Current state: 100% to stable
$ aws vpc-lattice get-rule --service-identifier svc-order \
  --listener-identifier listener-https --rule-identifier rule-default | \
  jq '.action.forward.targetGroups[] | {id: .targetGroupIdentifier, weight}'

{"id": "tg-stable", "weight": 100}
{"id": "tg-canary", "weight": 0}

# Shift 10% to canary
$ aws vpc-lattice update-rule \
  --service-identifier svc-order \
  --listener-identifier listener-https \
  --rule-identifier rule-default \
  --action '{
    "forward": {
      "targetGroups": [
        {"targetGroupIdentifier": "tg-stable", "weight": 90},
        {"targetGroupIdentifier": "tg-canary", "weight": 10}
      ]
    }
  }'

# Monitor canary error rate
$ watch -n 5 'aws cloudwatch get-metric-statistics \
  --namespace VPCLattice \
  --metric-name HTTPCode_Target_5XX_Count \
  --dimensions Name=TargetGroup,Value=tg-canary \
  --start-time $(date -u -d "5 minutes ago" +%Y-%m-%dT%H:%M:%S) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%S) \
  --period 60 --statistics Sum | jq ".Datapoints[0].Sum // 0"'

0  # No errors — safe to increase

# Progressive rollout: 10% → 25% → 50% → 100%
$ for weight in 25 50 100; do
  echo "Shifting to ${weight}% canary..."
  aws vpc-lattice update-rule \
    --service-identifier svc-order \
    --listener-identifier listener-https \
    --rule-identifier rule-default \
    --action "{
      \"forward\": {
        \"targetGroups\": [
          {\"targetGroupIdentifier\": \"tg-stable\", \"weight\": $((100-weight))},
          {\"targetGroupIdentifier\": \"tg-canary\", \"weight\": ${weight}}
        ]
      }
    }"
  sleep 300  # 5 min between steps
done

Performance Comparison: Measured Over 30 Days

Controlled A/B test running identical workloads through both meshes on the same cluster:

MetricApp Mesh (Envoy)VPC LatticeChange
P50 latency4.2ms1.8ms-57%
P95 latency12.1ms4.8ms-60%
P99 latency18.4ms6.2ms-66%
Connection establishment8.1ms2.4ms-70%
Memory per service85MB0MB-100%
CPU per service0.15 vCPU0 vCPU-100%
Container startup impact+2.4s+0s-100%
Cross-account latency12ms3.1ms-74%
Monthly incidents (networking)235-78%
MTTR (networking issues)45min12min-73%

Statistical significance: n=4.2M requests per mesh over 30 days, p < 0.001 for all latency comparisons using Welch's t-test.

Observability Without Sidecars

VPC Lattice logs to CloudWatch, but the depth differs from Envoy. Here's how we maintain observability:

# Enable access logs for the service network
$ aws vpc-lattice create-access-log-subscription \
  --resource-identifier sn-0a1b2c3d4e5f \
  --destination-arn arn:aws:logs:me-south-1:111122223333:log-group:/aws/vpc-lattice/platform-network

# Query logs with CloudWatch Insights
$ aws logs start-query \
  --log-group-name /aws/vpc-lattice/platform-network \
  --query-string '
    fields @timestamp, sourceVpcId, targetGroupArn, requestMethod, 
           requestPath, responseCode, duration
    | filter responseCode >= 500
    | stats count() as errors by targetGroupArn, bin(5m)
    | sort errors desc
    | limit 20
  ' \
  --start-time $(date -d '1 hour ago' +%s) \
  --end-time $(date +%s)

Sample access log entry:

{
  "timestamp": "2026-01-15T14:22:08.123Z",
  "serviceNetworkArn": "arn:aws:vpc-lattice:me-south-1:111122223333:servicenetwork/sn-0a1b",
  "serviceArn": "arn:aws:vpc-lattice:me-south-1:111122223333:service/svc-order",
  "sourceVpcId": "vpc-0abc123",
  "sourceIp": "10.0.1.45",
  "targetGroupArn": "arn:aws:vpc-lattice:me-south-1:111122223333:targetgroup/tg-stable",
  "targetIp": "10.0.2.67",
  "requestMethod": "POST",
  "requestPath": "/orders",
  "responseCode": 200,
  "duration": 23,
  "requestContentLength": 1247,
  "responseContentLength": 892,
  "callerPrincipal": "arn:aws:iam::111122223333:role/payment-service-role"
}

Limitations and Engineering Workarounds

VPC Lattice is not a full Envoy replacement. Here's what we lost and how we compensated:

CapabilityApp MeshVPC LatticeOur Workaround
Circuit breakingNative (Envoy)Not availableApplication-level with opossum
Rate limitingPer-routeNot availableAPI Gateway fronting + application-level
Custom retry policiesFull Envoy configBasic health checksExponential backoff in SDK wrapper
Request mirroringNativeNot availableCustom Lambda for shadow traffic
Distributed tracingAuto-instrumentedHeader propagation onlyAWS X-Ray SDK + OpenTelemetry
gRPC supportFullSupported (GA 2024)Works natively
WebSocketSupportedNot supportedKeep ALB for WebSocket services
// Our SDK wrapper adds retry + circuit breaker (replacing Envoy's native support)
import CircuitBreaker from 'opossum';

const latticeService = (serviceName: string) => {
  const baseUrl = `https://${serviceName}.platform-network.vpc-lattice-me-south-1.on.aws`;
  
  const call = async (path: string, options: RequestInit = {}) => {
    const response = await fetch(`${baseUrl}${path}`, {
      ...options,
      headers: {
        ...options.headers,
        'x-request-id': crypto.randomUUID(),
        'x-trace-id': getTraceId(),
      },
    });
    if (!response.ok) throw new Error(`${response.status}: ${response.statusText}`);
    return response.json();
  };

  return new CircuitBreaker(call, {
    timeout: 3000,
    errorThresholdPercentage: 50,
    resetTimeout: 30000,
    volumeThreshold: 10,
  });
};

// Usage
const orderService = latticeService('order-service');
const result = await orderService.fire('/orders/ORD-482917');

Migration Runbook Summary

Our migration took 6 weeks for 18 services. The process per service:

# Week 1-2: Set up VPC Lattice alongside existing App Mesh
$ terraform apply -target=module.vpc_lattice_service_network
$ terraform apply -target=module.vpc_lattice_services

# Week 3-4: Dual-run (both meshes active, VPC Lattice at 10%)
$ aws vpc-lattice update-rule ... --weight 10  # Canary via Lattice

# Week 5: Shift to 100% VPC Lattice
$ aws vpc-lattice update-rule ... --weight 100

# Week 6: Remove App Mesh infrastructure
$ terraform destroy -target=module.app_mesh
$ kubectl delete namespace appmesh-system

# Verify: No more sidecar containers
$ kubectl get pods --all-namespaces | grep envoy | wc -l
0

Cost Impact (Final Numbers)

ComponentBefore (App Mesh)After (VPC Lattice)Monthly Savings
Sidecar compute$4,200$0$4,200
VPC Lattice processing$0$1,200 (48TB × $0.025/GB)-$1,200
Engineering ops time$8,000$2,400$5,600
Cross-account networking$1,800$0$1,800
Total$14,000$3,600$10,400 (74% savings)

Annual savings: $124,800. Migration cost (6 weeks × 2 engineers): ~$48,000. Payback period: 4.6 months.

Key Takeaways

  1. VPC Lattice eliminates the sidecar tax: Zero memory, zero CPU, zero startup latency. For services behind ECS or EKS, this alone justifies the migration.
  2. Cross-account communication is the killer feature: No VPC peering, no Transit Gateway, no DNS gymnastics. Services in different accounts just talk — this simplified our multi-account architecture more than any other single change.
  3. IAM auth replaces mTLS complexity: Certificate rotation is AWS's problem. Policy changes are instant (no rolling deploy of sidecars). Integration with AWS Organizations makes org-wide policies trivial.
  4. You lose advanced traffic management: Circuit breaking, rate limiting, and request mirroring must move to application code. This is acceptable if your team treats resilience as an application concern (which it should be anyway).
  5. Weighted routing enables safe deployments: Native canary support without additional infrastructure. Combined with CloudWatch alarms, automated rollback becomes a 10-line script.

VPC Lattice is not App Mesh v2. It's a fundamentally different approach — simpler, faster, cheaper, but with fewer knobs. For 90% of service-to-service communication, fewer knobs is exactly what you want. The remaining 10% (WebSocket, request mirroring, complex retry) still needs purpose-built solutions.

Frequently Asked Questions

How does VPC Lattice pricing compare to App Mesh at scale?

App Mesh itself is free, but the sidecars consume compute. At our scale (340 sidecars on Fargate), sidecar cost was $4,200/month. VPC Lattice charges $0.025/GB of data processed. We process ~48TB/month, costing $1,200. Net savings: $3,000/month on compute alone, before counting engineering time savings.

Can VPC Lattice replace an API Gateway?

No. VPC Lattice handles east-west traffic (service-to-service). API Gateway handles north-south traffic (client-to-service) with features like throttling, API keys, request transformation, and WAF integration. We use both: API Gateway for external traffic, VPC Lattice for internal.

Does VPC Lattice work with EKS pod-level targeting?

Yes, since late 2024. You register pods as IP targets in a target group. The Kubernetes Gateway API controller for VPC Lattice automates target registration as pods scale up/down. We use this for our EKS workloads alongside ECS Fargate tasks in the same service network.

What's the maximum number of services per service network?

AWS documents 500 services per service network and 500 service network associations per VPC. At our 18-service scale, we're well within limits. Larger organizations may need multiple service networks organized by domain (e.g., payments-network, logistics-network).

How do you handle service discovery without Envoy's EDS?

VPC Lattice generates a DNS name per service (e.g., order-service.platform-network.vpc-lattice-me-south-1.on.aws). Services call this DNS name directly. Target registration (adding/removing IPs as tasks scale) is handled by AWS or the Gateway API controller — no Envoy EDS polling loop needed.

Comments

    No comments yet. Be the first to share your thoughts.