API Gateway Compression Strategies: Reducing Transfer Costs by 67%
Response compression at the API gateway layer reducing data transfer costs by 67% with minimal latency overhead.

API responses are larger than they need to be. JSON is verbose, headers are redundant, and most APIs send the same structural overhead with every response. After implementing compression at our API gateway layer, we reduced response sizes by 67% on average — translating directly to 67% lower data transfer costs and 40% faster response times for clients on constrained networks.
This guide covers compression strategy selection, implementation at the gateway layer, and the performance trade-offs you need to understand.
The Cost of Uncompressed APIs
Our API platform serves 45 million requests per day with an average response size of 4.2 KB. Monthly egress:
- 45M requests/day x 4.2 KB = 189 TB/month (raw)
- At $0.09/GB egress: $17,010/month in data transfer alone
After compression (average 67% reduction):
- 45M requests/day x 1.4 KB = 63 TB/month (compressed)
- At $0.09/GB egress: $5,670/month
- Monthly savings: $11,340
Compression is often the easiest win in a larger AWS data transfer cost optimization strategy that also covers VPC endpoints and cross-AZ routing.
Similar savings patterns appear when addressing NAT Gateway egress costs and Elastic IP hidden charges — small per-unit costs that compound into massive bills at scale.
Compression Algorithm Comparison
We benchmarked three compression algorithms against our actual API response corpus:
| Algorithm | Compression Ratio | Compression Speed | Decompression Speed | CPU Cost |
|---|---|---|---|---|
| gzip (level 6) | 62% reduction | 145 MB/s | 380 MB/s | Medium |
| Brotli (level 4) | 71% reduction | 98 MB/s | 420 MB/s | High |
| zstd (level 3) | 67% reduction | 310 MB/s | 890 MB/s | Low |
For API responses (typically 1-50 KB JSON), Brotli provides the best compression ratio but at significant CPU cost. We chose a hybrid approach: Brotli for responses > 10 KB, gzip for smaller responses. Pairing compression with a CDN origin shield further reduces the bandwidth that even needs compressing.
Implementation: API Gateway Middleware
import gzip
import zlib
from io import BytesIO
from typing import Optional, Callable
from dataclasses import dataclass
from functools import lru_cache
try:
import brotli
BROTLI_AVAILABLE = True
except ImportError:
BROTLI_AVAILABLE = False
@dataclass
class CompressionConfig:
min_size_bytes: int = 256 # Don't compress below this
max_size_bytes: int = 10_485_760 # 10 MB max
gzip_level: int = 6
brotli_quality: int = 4
brotli_threshold: int = 10_240 # Use brotli above 10 KB
excluded_content_types: tuple = (
'image/png', 'image/jpeg', 'image/webp',
'application/zip', 'application/gzip'
)
class CompressionMiddleware:
"""API Gateway compression with algorithm negotiation."""
def __init__(self, config: Optional[CompressionConfig] = None):
self.config = config or CompressionConfig()
self._stats = {
'total_requests': 0,
'compressed_requests': 0,
'bytes_before': 0,
'bytes_after': 0
}
def should_compress(self, content_type: str, content_length: int) -> bool:
"""Determine if response should be compressed."""
if content_length < self.config.min_size_bytes:
return False
if content_length > self.config.max_size_bytes:
return False
if content_type in self.config.excluded_content_types:
return False
return True
def negotiate_algorithm(
self,
accept_encoding: str,
content_length: int
) -> Optional[str]:
"""Select best compression algorithm based on client support and payload size."""
encodings = self._parse_accept_encoding(accept_encoding)
# Prefer brotli for large payloads if client supports it
if (BROTLI_AVAILABLE
and 'br' in encodings
and content_length >= self.config.brotli_threshold):
return 'br'
# Fall back to gzip (universally supported)
if 'gzip' in encodings:
return 'gzip'
# Deflate as last resort
if 'deflate' in encodings:
return 'deflate'
return None
def compress(self, data: bytes, algorithm: str) -> tuple[bytes, str]:
"""Compress data with specified algorithm. Returns (compressed_data, encoding)."""
self._stats['bytes_before'] += len(data)
if algorithm == 'br' and BROTLI_AVAILABLE:
compressed = brotli.compress(data, quality=self.config.brotli_quality)
elif algorithm == 'gzip':
compressed = gzip.compress(data, compresslevel=self.config.gzip_level)
elif algorithm == 'deflate':
compressed = zlib.compress(data, self.config.gzip_level)
else:
return data, 'identity'
self._stats['bytes_after'] += len(compressed)
self._stats['compressed_requests'] += 1
# Only use compressed version if it's actually smaller
if len(compressed) < len(data):
return compressed, algorithm
self._stats['bytes_after'] -= len(compressed)
self._stats['bytes_after'] += len(data)
return data, 'identity'
def _parse_accept_encoding(self, header: str) -> set[str]:
"""Parse Accept-Encoding header into set of supported algorithms."""
encodings = set()
for part in header.split(','):
encoding = part.strip().split(';')[0].strip().lower()
if encoding:
encodings.add(encoding)
return encodings
@property
def compression_ratio(self) -> float:
"""Overall compression ratio."""
if self._stats['bytes_before'] == 0:
return 0.0
return 1 - (self._stats['bytes_after'] / self._stats['bytes_before'])
Response Optimization: Beyond Compression
Compression alone doesn't address structural inefficiency. We layer three optimizations:
1. Field Filtering (Sparse Fieldsets)
from typing import Any
def filter_response_fields(
response: dict[str, Any],
fields: Optional[list[str]] = None,
exclude: Optional[list[str]] = None
) -> dict[str, Any]:
"""Filter response to only requested fields."""
if fields:
return {k: v for k, v in response.items() if k in fields}
if exclude:
return {k: v for k, v in response.items() if k not in exclude}
return response
# With ?fields=id,name,email — response drops from 4.2 KB to 0.3 KB
# That's a 93% reduction BEFORE compression
2. Response Envelope Optimization
import json
from typing import Any
class CompactSerializer:
"""Serialize API responses with minimal overhead."""
@staticmethod
def serialize_list(
items: list[dict],
total_count: int,
page: int,
page_size: int
) -> bytes:
"""Compact list serialization without redundant keys."""
# Instead of repeating keys for each object,
# use columnar format for homogeneous lists
if not items:
return json.dumps({'data': [], 'meta': {'total': total_count}}).encode()
keys = list(items[0].keys())
values = [[item.get(k) for k in keys] for item in items]
response = {
'_keys': keys,
'_values': values,
'meta': {
'total': total_count,
'page': page,
'pageSize': page_size
}
}
return json.dumps(response, separators=(',', ':')).encode()
Cost Impact Analysis
| Optimization Layer | Avg Response Size | Reduction | Monthly Transfer | Monthly Cost |
|---|---|---|---|---|
| No optimization | 4.2 KB | Baseline | 189 TB | $17,010 |
| gzip only | 1.6 KB | 62% | 72 TB | $6,480 |
| Brotli (large responses) | 1.4 KB | 67% | 63 TB | $5,670 |
| + Field filtering | 0.8 KB | 81% | 36 TB | $3,240 |
| + Compact serialization | 0.6 KB | 86% | 27 TB | $2,430 |
Performance Trade-offs
Compression is not free. CPU time spent compressing increases response latency:
| Algorithm | Compression Time (4 KB) | Added Latency | Break-even Network Speed |
|---|---|---|---|
| gzip-6 | 0.03 ms | Negligible | 1 Mbps |
| brotli-4 | 0.12 ms | Negligible | 5 Mbps |
| brotli-11 | 4.8 ms | Significant | Not worth it for APIs |
For API responses under 50 KB, compression latency is unmeasurable (< 1ms). The network time saved by sending fewer bytes always exceeds the compression CPU time unless the client is on a local gigabit network.
Gateway Configuration: AWS API Gateway
For AWS API Gateway, enable compression at the API level:
resource "aws_api_gateway_rest_api" "main" {
name = "content-api"
description = "Content delivery API with compression"
minimum_compression_size = 256 # Compress responses > 256 bytes
endpoint_configuration {
types = ["REGIONAL"]
}
}
Monitoring Compression Effectiveness
Key metrics to track:
- Compression ratio by endpoint: Identifies poorly-compressible responses
- CPU utilization on gateway nodes: Ensures compression isn't saturating compute
- Client-reported latency: Confirms network savings exceed compression time
- Cache hit ratio impact: Compressed responses may reduce cache effectiveness if varies by client
Key Takeaways
- 67% average reduction is achievable. JSON APIs compress extremely well due to key repetition and whitespace.
- Algorithm choice depends on payload size. gzip for < 10 KB, Brotli for > 10 KB.
- Layer optimizations compound. Compression + field filtering + compact serialization = 86% total reduction.
- CPU cost is negligible for APIs. Sub-millisecond compression time for typical API payloads.
- Don't compress already-compressed content. Images, videos, and pre-compressed archives get larger with double-compression.
Compression is the lowest-effort, highest-impact data transfer optimization available. If your API gateway isn't compressing responses, you're overpaying by 60-70% on every response sent. For a broader approach to reducing transfer costs, see the AWS data transfer cost optimization guide covering VPC endpoints, cross-AZ routing, and regional data locality.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.