Cloud NAT Cost Traps and How We Reduced Egress Charges by 45%
A deep dive into GCP Cloud NAT pricing pitfalls, hidden egress costs, and the 5 optimization strategies that cut our networking bill by 45% at scale.

Google Cloud NAT is one of those services that seems simple until you get the bill. We were running 400+ GKE pods across three clusters, all routing outbound traffic through Cloud NAT. Our monthly networking bill hit $34,000 — and Cloud NAT-related charges accounted for $18,000 of that. After a focused optimization effort, we reduced it to $9,900. Here's how.
Understanding Cloud NAT Pricing
Cloud NAT pricing is deceptively complex. There are three cost dimensions that most teams miss:
| Cost Component | Rate | What Triggers It |
|---|---|---|
| NAT gateway processing | $0.045/GB | All outbound traffic through NAT |
| Port allocation | $0.006/hour per 64 ports | Each VM/pod gets allocated ports |
| Egress to internet | $0.08-0.12/GB | Standard GCP egress pricing applies ON TOP |
The critical insight: Cloud NAT processing fees are additive to standard egress charges. You pay both. A 1GB transfer to the internet costs $0.045 (NAT) + $0.085 (egress) = $0.13/GB instead of just the $0.085 egress.
The Port Allocation Trap
This was our biggest surprise. Cloud NAT allocates ports to each VM (or GKE node) in blocks of 64. If you have a node pool with 50 nodes, each with the default minimum of 64 ports allocated, you're paying:
50 nodes × 64 ports × $0.006/hour / 64 = $0.30/hour = $219/month
But here's the trap: if any single pod on a node needs more than 64 concurrent connections to a single destination IP, Cloud NAT automatically doubles the allocation. With our microservices architecture making hundreds of connections to external APIs, many nodes had 1024+ ports allocated.
50 nodes × 1024 ports × $0.006/hour / 64 = $4.80/hour = $3,504/month
Just from port allocation — before any data transfer.
Our Initial Architecture
The original setup was straightforward:
# Terraform - Cloud NAT configuration (before optimization)
resource "google_compute_router_nat" "main" {
name = "main-nat"
router = google_compute_router.main.name
region = "us-central1"
nat_ip_allocate_option = "AUTO_ONLY"
source_subnetwork_ip_ranges_to_nat = "ALL_SUBNETWORKS_ALL_IP_RANGES"
log_config {
enable = true
filter = "ALL"
}
# Default port allocation - this was the problem
min_ports_per_vm = 64
enable_dynamic_port_allocation = false
}
Every pod, regardless of whether it needed internet access, routed through NAT. Logging was capturing every connection. Port allocation was static and over-provisioned.
Optimization 1: Selective NAT with Subnet Segmentation
Not every workload needs outbound internet access. We segmented our VPC into NAT and non-NAT subnets:
resource "google_compute_router_nat" "optimized" {
name = "optimized-nat"
router = google_compute_router.main.name
region = "us-central1"
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.nat_ip.self_link]
# Only NAT specific subnets
source_subnetwork_ip_ranges_to_nat = "LIST_OF_SUBNETWORKS"
subnetwork {
name = google_compute_subnetwork.internet_facing.id
source_ip_ranges_to_nat = ["ALL_IP_RANGES"]
}
# Don't NAT internal-only workloads
# subnetwork for internal services is excluded
}
Impact: 35% of our pods (internal services, databases, caches) never needed internet access. Excluding them reduced NAT processing by $2,100/month.
Optimization 2: Dynamic Port Allocation
GCP introduced dynamic port allocation (DPA) which automatically adjusts port allocation per VM based on actual usage:
resource "google_compute_router_nat" "optimized" {
# ... other config ...
enable_dynamic_port_allocation = true
min_ports_per_vm = 32 # Start low
max_ports_per_vm = 4096 # Scale up if needed
# Enable endpoint-independent mapping for connection reuse
enable_endpoint_independent_mapping = false
}
With DPA enabled:
- Idle nodes dropped from 1024 ports to 32 ports
- Active nodes scaled up only when needed
- Overall port allocation dropped by 70%
Impact: Port allocation costs went from $3,504/month to $1,050/month.
Optimization 3: Private Google Access for GCP Services
A shocking amount of our "internet" traffic was actually going to other Google services — Cloud Storage, BigQuery, Container Registry, Pub/Sub. All of this was routing through NAT unnecessarily.
resource "google_compute_subnetwork" "main" {
name = "main-subnet"
ip_cidr_range = "10.0.0.0/20"
region = "us-central1"
network = google_compute_network.vpc.id
# Enable Private Google Access - traffic to Google APIs
# bypasses NAT entirely
private_ip_google_access = true
}
We also configured Private Service Connect for our most traffic-heavy Google API endpoints:
resource "google_compute_global_address" "private_service_connect" {
name = "psc-googleapis"
purpose = "PRIVATE_SERVICE_CONNECT"
address_type = "INTERNAL"
network = google_compute_network.vpc.id
address = "10.255.255.254"
}
Impact: 40% of our outbound "internet" traffic was actually to Google APIs. Redirecting it saved $3,600/month in NAT processing fees.
Optimization 4: Connection Draining and Timeout Tuning
Cloud NAT holds port mappings for idle connections based on timeout settings. Default timeouts are generous:
| Connection Type | Default Timeout | Our Optimized Value |
|---|---|---|
| TCP established | 1200s (20min) | 300s (5min) |
| TCP transitory | 30s | 15s |
| UDP | 30s | 15s |
| ICMP | 30s | 15s |
resource "google_compute_router_nat" "optimized" {
# ... other config ...
tcp_established_idle_timeout_sec = 300
tcp_transitory_idle_timeout_sec = 15
udp_idle_timeout_sec = 15
icmp_idle_timeout_sec = 15
}
Shorter timeouts mean ports get recycled faster, reducing the peak port allocation. Combined with DPA, this further reduced costs.
Impact: Additional $800/month reduction from faster port recycling.
Optimization 5: Egress Routing Through Internal Load Balancers
For traffic going to specific known destinations (partner APIs, SaaS webhooks), we deployed a fleet of proxy instances in the internet-facing subnet instead of routing all traffic through Cloud NAT:
# Envoy proxy pool for high-volume external API calls
resource "google_compute_instance_group_manager" "egress_proxy" {
name = "egress-proxy-pool"
base_instance_name = "egress-proxy"
zone = "us-central1-a"
target_size = 3
version {
instance_template = google_compute_instance_template.egress_proxy.id
}
auto_healing_policies {
health_check = google_compute_health_check.proxy_health.id
initial_delay_sec = 60
}
}
This approach gave us:
- Connection pooling (reducing total outbound connections by 80%)
- Request coalescing for burst traffic
- Better observability on external API calls
- Fixed egress IPs for allowlisting
Impact: $1,500/month saved from reduced NAT processing and better connection efficiency.
Results: Before and After
| Component | Before | After | Savings |
|---|---|---|---|
| NAT processing fees | $8,400 | $3,200 | $5,200 |
| Port allocation | $3,504 | $1,050 | $2,454 |
| Logging (reduced to errors only) | $1,200 | $200 | $1,000 |
| Egress (unchanged but rerouted) | $4,896 | $5,450 | -$554* |
| Total NAT-related costs | $18,000 | $9,900 | $8,100 (45%) |
*Egress increased slightly because Private Service Connect uses a different billing path for some traffic patterns.
Monitoring: Catching Cost Drift
We set up Cloud Monitoring alerts to catch cost regression:
# Alert policy for NAT port exhaustion
resource "google_monitoring_alert_policy" "nat_port_usage" {
display_name = "Cloud NAT Port Usage > 80%"
combiner = "OR"
conditions {
display_name = "NAT port allocation high"
condition_threshold {
filter = "resource.type=\"nat_gateway\" AND metric.type=\"router.googleapis.com/nat/allocated_ports\""
comparison = "COMPARISON_GT"
threshold_value = 0.8
duration = "300s"
aggregations {
alignment_period = "60s"
per_series_aligner = "ALIGN_MEAN"
}
}
}
notification_channels = [google_monitoring_notification_channel.ops_team.id]
}
Common Pitfalls to Avoid
-
Logging everything: Cloud NAT logs are verbose. At our scale, full logging added $1,200/month in Cloud Logging costs alone. Log only errors and dropped connections.
-
Auto-allocated IPs: When NAT auto-allocates IPs, you can't predict them for allowlisting. Use manual IP allocation.
-
Not monitoring port exhaustion: When ports run out, connections drop silently. Your app sees timeouts but Cloud NAT doesn't surface this clearly.
-
Ignoring IPv6: If your workloads can use IPv6, they bypass NAT entirely. GKE supports dual-stack networking.
Key Takeaways
-
Cloud NAT processing fees are additive — they stack on top of standard egress charges. Budget for both.
-
Enable Private Google Access immediately — there's no reason to route GCP-to-GCP traffic through NAT. This is free money.
-
Dynamic port allocation is essential at scale — static allocation over-provisions dramatically for bursty workloads.
-
Segment your subnets by internet-access need — not every service needs NAT. Excluding internal-only services is the easiest win.
-
Proxy pools beat NAT for high-volume external APIs — connection pooling and request coalescing reduce total connections and give better observability.
-
Monitor port allocation continuously — exhaustion causes silent failures that are hard to debug after the fact.
Cloud NAT is a necessary service for private workloads, but it requires active cost management. The default configuration optimizes for simplicity, not cost. Take 30 minutes to review your NAT configuration against these optimizations — the savings compound monthly.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.