Cloud NAT Cost Traps and How We Reduced Egress Charges by 45%

A deep dive into GCP Cloud NAT pricing pitfalls, hidden egress costs, and the 5 optimization strategies that cut our networking bill by 45% at scale.

#gcp#cloud-nat#cost-optimization#networking#egress
Cover image for the article: Cloud NAT Cost Traps and How We Reduced Egress Charges by 45%

Google Cloud NAT is one of those services that seems simple until you get the bill. We were running 400+ GKE pods across three clusters, all routing outbound traffic through Cloud NAT. Our monthly networking bill hit $34,000 — and Cloud NAT-related charges accounted for $18,000 of that. After a focused optimization effort, we reduced it to $9,900. Here's how.

Understanding Cloud NAT Pricing

Cloud NAT pricing is deceptively complex. There are three cost dimensions that most teams miss:

Cost ComponentRateWhat Triggers It
NAT gateway processing$0.045/GBAll outbound traffic through NAT
Port allocation$0.006/hour per 64 portsEach VM/pod gets allocated ports
Egress to internet$0.08-0.12/GBStandard GCP egress pricing applies ON TOP

The critical insight: Cloud NAT processing fees are additive to standard egress charges. You pay both. A 1GB transfer to the internet costs $0.045 (NAT) + $0.085 (egress) = $0.13/GB instead of just the $0.085 egress.

The Port Allocation Trap

This was our biggest surprise. Cloud NAT allocates ports to each VM (or GKE node) in blocks of 64. If you have a node pool with 50 nodes, each with the default minimum of 64 ports allocated, you're paying:

50 nodes × 64 ports × $0.006/hour / 64 = $0.30/hour = $219/month

But here's the trap: if any single pod on a node needs more than 64 concurrent connections to a single destination IP, Cloud NAT automatically doubles the allocation. With our microservices architecture making hundreds of connections to external APIs, many nodes had 1024+ ports allocated.

50 nodes × 1024 ports × $0.006/hour / 64 = $4.80/hour = $3,504/month

Just from port allocation — before any data transfer.

Our Initial Architecture

Cloud NAT Before Optimization

The original setup was straightforward:

# Terraform - Cloud NAT configuration (before optimization)
resource "google_compute_router_nat" "main" {
  name                               = "main-nat"
  router                             = google_compute_router.main.name
  region                             = "us-central1"
  nat_ip_allocate_option            = "AUTO_ONLY"
  source_subnetwork_ip_ranges_to_nat = "ALL_SUBNETWORKS_ALL_IP_RANGES"

  log_config {
    enable = true
    filter = "ALL"
  }

  # Default port allocation - this was the problem
  min_ports_per_vm = 64
  enable_dynamic_port_allocation = false
}

Every pod, regardless of whether it needed internet access, routed through NAT. Logging was capturing every connection. Port allocation was static and over-provisioned.

Optimization 1: Selective NAT with Subnet Segmentation

Not every workload needs outbound internet access. We segmented our VPC into NAT and non-NAT subnets:

resource "google_compute_router_nat" "optimized" {
  name                               = "optimized-nat"
  router                             = google_compute_router.main.name
  region                             = "us-central1"
  nat_ip_allocate_option            = "MANUAL_ONLY"
  nat_ips                            = [google_compute_address.nat_ip.self_link]

  # Only NAT specific subnets
  source_subnetwork_ip_ranges_to_nat = "LIST_OF_SUBNETWORKS"

  subnetwork {
    name                    = google_compute_subnetwork.internet_facing.id
    source_ip_ranges_to_nat = ["ALL_IP_RANGES"]
  }

  # Don't NAT internal-only workloads
  # subnetwork for internal services is excluded
}

Impact: 35% of our pods (internal services, databases, caches) never needed internet access. Excluding them reduced NAT processing by $2,100/month.

Optimization 2: Dynamic Port Allocation

GCP introduced dynamic port allocation (DPA) which automatically adjusts port allocation per VM based on actual usage:

resource "google_compute_router_nat" "optimized" {
  # ... other config ...

  enable_dynamic_port_allocation = true
  min_ports_per_vm              = 32    # Start low
  max_ports_per_vm              = 4096  # Scale up if needed

  # Enable endpoint-independent mapping for connection reuse
  enable_endpoint_independent_mapping = false
}

With DPA enabled:

  • Idle nodes dropped from 1024 ports to 32 ports
  • Active nodes scaled up only when needed
  • Overall port allocation dropped by 70%

Impact: Port allocation costs went from $3,504/month to $1,050/month.

Optimization 3: Private Google Access for GCP Services

A shocking amount of our "internet" traffic was actually going to other Google services — Cloud Storage, BigQuery, Container Registry, Pub/Sub. All of this was routing through NAT unnecessarily.

resource "google_compute_subnetwork" "main" {
  name          = "main-subnet"
  ip_cidr_range = "10.0.0.0/20"
  region        = "us-central1"
  network       = google_compute_network.vpc.id

  # Enable Private Google Access - traffic to Google APIs
  # bypasses NAT entirely
  private_ip_google_access = true
}

We also configured Private Service Connect for our most traffic-heavy Google API endpoints:

resource "google_compute_global_address" "private_service_connect" {
  name         = "psc-googleapis"
  purpose      = "PRIVATE_SERVICE_CONNECT"
  address_type = "INTERNAL"
  network      = google_compute_network.vpc.id
  address      = "10.255.255.254"
}

Impact: 40% of our outbound "internet" traffic was actually to Google APIs. Redirecting it saved $3,600/month in NAT processing fees.

Optimization 4: Connection Draining and Timeout Tuning

Cloud NAT holds port mappings for idle connections based on timeout settings. Default timeouts are generous:

Connection TypeDefault TimeoutOur Optimized Value
TCP established1200s (20min)300s (5min)
TCP transitory30s15s
UDP30s15s
ICMP30s15s
resource "google_compute_router_nat" "optimized" {
  # ... other config ...

  tcp_established_idle_timeout_sec = 300
  tcp_transitory_idle_timeout_sec  = 15
  udp_idle_timeout_sec            = 15
  icmp_idle_timeout_sec           = 15
}

Shorter timeouts mean ports get recycled faster, reducing the peak port allocation. Combined with DPA, this further reduced costs.

Impact: Additional $800/month reduction from faster port recycling.

Optimization 5: Egress Routing Through Internal Load Balancers

For traffic going to specific known destinations (partner APIs, SaaS webhooks), we deployed a fleet of proxy instances in the internet-facing subnet instead of routing all traffic through Cloud NAT:

# Envoy proxy pool for high-volume external API calls
resource "google_compute_instance_group_manager" "egress_proxy" {
  name               = "egress-proxy-pool"
  base_instance_name = "egress-proxy"
  zone               = "us-central1-a"
  target_size        = 3

  version {
    instance_template = google_compute_instance_template.egress_proxy.id
  }

  auto_healing_policies {
    health_check      = google_compute_health_check.proxy_health.id
    initial_delay_sec = 60
  }
}

This approach gave us:

  • Connection pooling (reducing total outbound connections by 80%)
  • Request coalescing for burst traffic
  • Better observability on external API calls
  • Fixed egress IPs for allowlisting

Impact: $1,500/month saved from reduced NAT processing and better connection efficiency.

Results: Before and After

Cloud NAT Cost Reduction

ComponentBeforeAfterSavings
NAT processing fees$8,400$3,200$5,200
Port allocation$3,504$1,050$2,454
Logging (reduced to errors only)$1,200$200$1,000
Egress (unchanged but rerouted)$4,896$5,450-$554*
Total NAT-related costs$18,000$9,900$8,100 (45%)

*Egress increased slightly because Private Service Connect uses a different billing path for some traffic patterns.

Monitoring: Catching Cost Drift

We set up Cloud Monitoring alerts to catch cost regression:

# Alert policy for NAT port exhaustion
resource "google_monitoring_alert_policy" "nat_port_usage" {
  display_name = "Cloud NAT Port Usage > 80%"
  combiner     = "OR"

  conditions {
    display_name = "NAT port allocation high"
    condition_threshold {
      filter          = "resource.type=\"nat_gateway\" AND metric.type=\"router.googleapis.com/nat/allocated_ports\""
      comparison      = "COMPARISON_GT"
      threshold_value = 0.8
      duration        = "300s"
      aggregations {
        alignment_period   = "60s"
        per_series_aligner = "ALIGN_MEAN"
      }
    }
  }

  notification_channels = [google_monitoring_notification_channel.ops_team.id]
}

Common Pitfalls to Avoid

  1. Logging everything: Cloud NAT logs are verbose. At our scale, full logging added $1,200/month in Cloud Logging costs alone. Log only errors and dropped connections.

  2. Auto-allocated IPs: When NAT auto-allocates IPs, you can't predict them for allowlisting. Use manual IP allocation.

  3. Not monitoring port exhaustion: When ports run out, connections drop silently. Your app sees timeouts but Cloud NAT doesn't surface this clearly.

  4. Ignoring IPv6: If your workloads can use IPv6, they bypass NAT entirely. GKE supports dual-stack networking.

Key Takeaways

  1. Cloud NAT processing fees are additive — they stack on top of standard egress charges. Budget for both.

  2. Enable Private Google Access immediately — there's no reason to route GCP-to-GCP traffic through NAT. This is free money.

  3. Dynamic port allocation is essential at scale — static allocation over-provisions dramatically for bursty workloads.

  4. Segment your subnets by internet-access need — not every service needs NAT. Excluding internal-only services is the easiest win.

  5. Proxy pools beat NAT for high-volume external APIs — connection pooling and request coalescing reduce total connections and give better observability.

  6. Monitor port allocation continuously — exhaustion causes silent failures that are hard to debug after the fact.

Cloud NAT is a necessary service for private workloads, but it requires active cost management. The default configuration optimizes for simplicity, not cost. Take 30 minutes to review your NAT configuration against these optimizations — the savings compound monthly.

Comments

    No comments yet. Be the first to share your thoughts.