Exposing SaaS Services Securely with AWS PrivateLink: Architecture and Cost Model
A deep dive into AWS PrivateLink architecture for SaaS providers, covering endpoint services, cost modeling, and security patterns for private connectivity.

When you run a SaaS platform, every customer eventually asks the same question: "Can we connect without going over the public internet?" The answer used to involve VPN tunnels, peering agreements, and weeks of networking configuration. AWS PrivateLink changed that equation entirely.
Over the past two years, I have architected PrivateLink-based service exposure for three different B2B SaaS products. The pattern is elegant, but the cost model catches teams off guard. This article covers the full architecture, real-world cost data, and the operational decisions that determine whether PrivateLink becomes a competitive advantage or a budget sinkhole.
AWS PrivateLink Pricing
AWS PrivateLink charges $0.01 per hour per VPC endpoint ($7.20/month) plus $0.01 per GB of data processed. For a service processing 100 GB/month through a single endpoint, the total cost is approximately $8.20/month. There are no additional charges for the service provider side beyond the Network Load Balancer costs.
Why PrivateLink for SaaS Service Exposure
Traditional approaches to private connectivity between SaaS providers and enterprise customers involve VPC peering or site-to-site VPNs. Both create operational overhead that scales poorly.
| Approach | Setup Time | IP Overlap Handling | Ongoing Ops | Security Boundary |
|---|---|---|---|---|
| VPC Peering | 2-4 hours | None (breaks) | Medium | Shared routing |
| Site-to-Site VPN | 1-2 weeks | NAT required | High | Tunnel encryption |
| AWS PrivateLink | 30-60 min | Native support | Low | Unidirectional |
| Transit Gateway | 4-8 hours | Route table mgmt | Medium | Centralized |
PrivateLink wins because it is unidirectional. The customer initiates the connection; you never gain access to their VPC. This asymmetry eliminates entire classes of security concerns that enterprise procurement teams fixate on during evaluations.
Architecture: Provider Side
The provider architecture requires a Network Load Balancer (NLB) fronting your service, an Endpoint Service configuration, and a principal allowlist.
# Terraform: PrivateLink Endpoint Service
resource "aws_vpc_endpoint_service" "saas_api" {
acceptance_required = true
network_load_balancer_arns = [aws_lb.saas_nlb.arn]
allowed_principals = var.customer_account_arns
tags = {
Name = "saas-api-endpoint-service"
Environment = "production"
CostCenter = "platform-networking"
}
}
resource "aws_lb" "saas_nlb" {
name = "saas-api-nlb"
internal = true
load_balancer_type = "network"
subnets = var.private_subnet_ids
enable_cross_zone_load_balancing = true
}
resource "aws_lb_target_group" "saas_api" {
name = "saas-api-tg"
port = 443
protocol = "TLS"
vpc_id = var.vpc_id
health_check {
protocol = "TCP"
interval = 10
healthy_threshold = 3
unhealthy_threshold = 3
}
}
The critical architectural decision is whether to use a single NLB with multiple target groups or separate NLBs per customer tier. I recommend separate NLBs when you need per-customer rate limiting at the network level and shared NLBs when your application layer handles tenant isolation.
Architecture: Consumer Side
On the customer side, the setup is straightforward. They create an interface VPC endpoint pointing to your service name and receive a private DNS entry within their VPC.
# Customer-side Terraform
resource "aws_vpc_endpoint" "saas_provider" {
vpc_id = var.customer_vpc_id
service_name = "com.amazonaws.vpce.us-east-1.vpce-svc-0a1b2c3d4e5f6g7h8"
vpc_endpoint_type = "Interface"
subnet_ids = var.private_subnet_ids
security_group_ids = [aws_security_group.saas_endpoint.id]
private_dns_enabled = false
}
resource "aws_route53_zone" "private_saas" {
name = "api.saasprovider.internal"
vpc {
vpc_id = var.customer_vpc_id
}
}
resource "aws_route53_record" "saas_endpoint" {
zone_id = aws_route53_zone.private_saas.zone_id
name = "api.saasprovider.internal"
type = "A"
alias {
name = aws_vpc_endpoint.saas_provider.dns_entry[0]["dns_name"]
zone_id = aws_vpc_endpoint.saas_provider.dns_entry[0]["hosted_zone_id"]
evaluate_target_health = true
}
}
The Cost Model: What Nobody Tells You
PrivateLink pricing has two components that compound faster than most teams expect:
- Endpoint hourly cost: $0.01/hour per AZ per endpoint ($7.30/month per AZ)
- Data processing: $0.01/GB for the first 1 PB
For a typical B2B SaaS with 50 enterprise customers, each using 3 AZs for high availability:
| Customers | AZs | Monthly Endpoint Cost | Data (TB/mo) | Processing Cost | Total Monthly |
|---|---|---|---|---|---|
| 10 | 3 | $219 | 2 | $20 | $239 |
| 50 | 3 | $1,095 | 15 | $150 | $1,245 |
| 200 | 3 | $4,380 | 80 | $800 | $5,180 |
| 500 | 3 | $10,950 | 250 | $2,500 | $13,450 |
At 500 customers, you are spending over $160K annually on PrivateLink alone. This is where the pricing strategy becomes an engineering decision: do you absorb this as platform cost, charge per-connection fees, or bundle it into enterprise tier pricing?
Multi-Region PrivateLink: The Complexity Multiplier
When customers demand multi-region connectivity, costs multiply by region count. But the architecture also introduces latency considerations. Cross-region PrivateLink is not natively supported; you need an NLB per region with your service running in each.
# Architecture decision record
decision: Multi-region PrivateLink topology
status: accepted
context: |
Enterprise customers require <50ms latency from eu-west-1 and us-east-1.
Single-region PrivateLink adds 80-120ms cross-Atlantic latency.
options:
- Deploy NLB + endpoint service per region (higher cost, lower latency)
- Single region with Global Accelerator (lower cost, higher latency)
- Regional NLB with cross-region backend routing (medium cost, complex)
chosen: Deploy NLB + endpoint service per region
rationale: |
Latency requirements eliminate single-region options.
Cross-region backend routing adds failure modes without cost savings.
Security Patterns That Matter
Beyond the basic architecture, three security patterns separate production-grade PrivateLink deployments from proof-of-concept setups:
1. Principal-Based Access Control
resource "aws_vpc_endpoint_service_allowed_principal" "customer" {
for_each = toset(var.approved_customer_arns)
vpc_endpoint_service_id = aws_vpc_endpoint_service.saas_api.id
principal_arn = each.value
}
2. Connection Acceptance Automation
Manual acceptance does not scale past 20 customers. Automate with EventBridge:
import boto3
import json
def handle_connection_request(event, context):
ec2 = boto3.client('ec2')
detail = event['detail']
endpoint_id = detail['vpc-endpoint-id']
service_id = detail['service-id']
# Validate against customer database
customer = lookup_customer_by_aws_account(detail['owner'])
if customer and customer.tier in ['enterprise', 'premium']:
ec2.accept_vpc_endpoint_connections(
ServiceId=service_id,
VpcEndpointIds=[endpoint_id]
)
notify_customer_success(customer, endpoint_id)
else:
ec2.reject_vpc_endpoint_connections(
ServiceId=service_id,
VpcEndpointIds=[endpoint_id]
)
notify_sales_team(detail['owner'])
3. Per-Customer Traffic Monitoring
Tag NLB target groups per customer to get CloudWatch metrics at the tenant level. This enables per-customer SLA reporting and anomaly detection.
Operational Lessons from Production
After running PrivateLink at scale for 18 months, these are the operational realities:
Health check propagation delay: When your backend goes unhealthy, NLB health checks take 30-90 seconds to deregister targets. During this window, PrivateLink connections receive TCP RST packets. Customers need retry logic in their clients.
Cross-account visibility gap: You cannot see the customer endpoint health from the provider side. Build a heartbeat mechanism where customers periodically call a /health endpoint and you track response times.
AZ mapping mismatch: AWS AZ IDs (use1-az1) differ from AZ names (us-east-1a) across accounts. Always communicate using AZ IDs to ensure customers connect to the correct physical zone.
When Not to Use PrivateLink
PrivateLink is wrong for:
- High-throughput streaming: The $0.01/GB processing fee makes it expensive for data-intensive workloads above 100 TB/month. Use VPC peering with Transit Gateway instead.
- Bidirectional communication: If your service needs to call back into the customer VPC, PrivateLink cannot help. You need a reverse connection or mutual PrivateLink endpoints.
- Fewer than 5 enterprise customers: The engineering overhead of maintaining the endpoint service, acceptance automation, and monitoring does not justify the benefit below this threshold.
Key Takeaways
- PrivateLink is a product feature, not just infrastructure. Enterprise customers pay premium pricing partly for private connectivity. Price it into your tiers accordingly.
- Cost scales with customer count times AZ count. Model this early. At 200+ customers across 3 AZs, you are looking at $50K+ annually before data processing fees.
- Automate connection acceptance from day one. Manual approval creates a support bottleneck that slows enterprise onboarding.
- Plan for multi-region from the start. Retrofitting regional endpoint services is significantly harder than deploying them alongside your initial architecture.
- Monitor the consumer side indirectly. Build heartbeat and health check systems since you have zero visibility into customer endpoint status.
The best PrivateLink implementations I have seen treat the endpoint service as a first-class product surface: versioned, monitored, documented, and priced deliberately. The worst treat it as a networking checkbox and discover the cost and operational implications after 50 customers are connected.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.