Reducing CloudWatch Logs Costs from $12K to $3.2K/Month Without Losing Visibility
A systematic approach to log aggregation cost optimization using tiered retention, sampling, and structured logging that cut our bill by 73%

Our CloudWatch Logs bill hit $12,400/month and was growing 15% month-over-month. The root cause was not excessive infrastructure but undisciplined logging: verbose debug logs in production, duplicated log streams, and uniform 30-day retention across all log groups regardless of value. We implemented a tiered logging strategy that reduced costs to $3,200/month while actually improving our ability to diagnose production issues.
The Problem: All Logs Treated Equal
Every log line has a cost. At $0.50 per GB ingested and $0.03 per GB stored, CloudWatch Logs becomes expensive quickly when you are ingesting 820 GB per month without discrimination.
Cost breakdown before optimization:
- Ingestion (820 GB x $0.50): $410/month
- Storage (24.6 TB x $0.03): $738/month
- But wait: we had cross-region replication enabled on all log groups
- Total with replication and data transfer: $12,400/month
The problem was cultural as much as technical. Engineers added console.log liberally during development and never removed it. Health check endpoints logged every request. HTTP response bodies were logged at INFO level. Nobody owned log hygiene.
Architecture: Tiered Log Strategy
We classified all logs into four tiers based on their debugging value and compliance requirements:
| Tier | Description | Retention | Storage | Example |
|---|---|---|---|---|
| Critical | Security, compliance, errors | 365 days | CloudWatch + S3 | Auth failures, 5xx errors |
| Standard | Application business logic | 30 days | CloudWatch | Request logs, state transitions |
| Debug | Development diagnostics | 3 days | CloudWatch | Verbose debugging, SQL queries |
| Ephemeral | Health checks, heartbeats | 1 day | Filtered out pre-ingestion | /health, /ready, keep-alives |
Step 1: Structured Logging Standard
The foundation of cost optimization is structured logging. Unstructured text logs cannot be filtered efficiently. We deployed a shared logging library across all services:
import pino from 'pino';
interface LogConfig {
service: string;
environment: string;
tier?: 'critical' | 'standard' | 'debug' | 'ephemeral';
}
export function createLogger(config: LogConfig) {
return pino({
level: process.env.LOG_LEVEL || 'info',
formatters: {
level(label: string) {
return { level: label };
},
},
base: {
service: config.service,
environment: config.environment,
tier: config.tier || 'standard',
version: process.env.APP_VERSION,
},
// Redact sensitive fields automatically
redact: {
paths: [
'req.headers.authorization',
'req.headers.cookie',
'body.password',
'body.token',
'body.creditCard',
],
censor: '[REDACTED]',
},
// Custom serializers for cost-efficient logging
serializers: {
req: (req) => ({
method: req.method,
url: req.url,
// Don't log full headers - saves ~200 bytes per request
contentLength: req.headers['content-length'],
userAgent: req.headers['user-agent'],
}),
res: (res) => ({
statusCode: res.statusCode,
// Don't log response bodies at info level
contentLength: res.headers?.['content-length'],
}),
},
});
}
// Usage
const logger = createLogger({ service: 'order-service', environment: 'production' });
// Tier-aware logging
logger.info({ tier: 'standard', orderId, action: 'created' }, 'Order created');
logger.error({ tier: 'critical', orderId, error: err.message }, 'Payment failed');
logger.debug({ tier: 'debug', query, params, duration }, 'Database query executed');
Structured logging with tier annotations enabled automated routing: critical logs go to long-term storage, debug logs get short retention, and ephemeral logs get dropped before ingestion.
Step 2: Pre-Ingestion Filtering with Subscription Filters
The biggest cost win was preventing low-value logs from ever reaching CloudWatch. We use CloudWatch subscription filters to route logs before they incur ingestion charges:
import boto3
def configure_log_group_filters(log_group_name: str, service_config: dict):
"""Configure subscription filters for cost-optimized log routing."""
logs = boto3.client('logs')
# Filter out health check logs entirely (saves ~30% of volume)
logs.put_subscription_filter(
logGroupName=log_group_name,
filterName='drop-health-checks',
filterPattern='{ $.url != "/health" && $.url != "/ready" && $.url != "/metrics" }',
destinationArn=service_config['firehose_arn'],
roleArn=service_config['role_arn'],
)
# Set retention based on tier
retention_map = {
'critical': 365,
'standard': 30,
'debug': 3,
'ephemeral': 1,
}
tier = service_config.get('tier', 'standard')
logs.put_retention_policy(
logGroupName=log_group_name,
retentionInDays=retention_map[tier],
)
# Enable S3 archival for critical logs
if tier == 'critical':
logs.put_subscription_filter(
logGroupName=log_group_name,
filterName='archive-to-s3',
filterPattern='{ $.tier = "critical" }',
destinationArn=service_config['s3_firehose_arn'],
roleArn=service_config['role_arn'],
)
Health check filtering alone saved 30% of our ingestion volume. These endpoints are called every 10 seconds by load balancers and Kubernetes probes, generating massive log volume with zero debugging value.
Step 3: Log Sampling for High-Volume Endpoints
Some endpoints generate legitimate logs at volumes that are cost-prohibitive to retain at 100%. We implemented probabilistic sampling for successful requests while keeping 100% of errors:
import { createLogger } from './logger';
const logger = createLogger({ service: 'api-gateway', environment: 'production' });
function shouldSample(statusCode: number, path: string): boolean {
// Always log errors
if (statusCode >= 400) return true;
// Always log slow requests
// (checked elsewhere via response time)
// Sample high-volume endpoints at 10%
const highVolumeEndpoints = ['/api/v1/products', '/api/v1/search', '/api/v1/recommendations'];
if (highVolumeEndpoints.some(ep => path.startsWith(ep))) {
return Math.random() < 0.1;
}
// Sample all other successful requests at 25%
return Math.random() < 0.25;
}
export function requestLogger(req: Request, res: Response, next: () => void) {
const start = Date.now();
res.on('finish', () => {
const duration = Date.now() - start;
const shouldLog = shouldSample(res.statusCode, req.path) || duration > 2000;
if (shouldLog) {
logger.info({
tier: res.statusCode >= 500 ? 'critical' : 'standard',
method: req.method,
path: req.path,
statusCode: res.statusCode,
duration,
sampled: true,
sampleRate: getSampleRate(req.path),
}, 'Request completed');
}
});
next();
}
Sampling at 10-25% for successful requests provides sufficient statistical data for dashboards and trend analysis while dramatically reducing volume.
Step 4: Cross-Region Replication Elimination
Our largest unnecessary cost was cross-region log replication. Every log group was replicated to a DR region "just in case." Analysis showed that DR log access happened zero times in 18 months. We replaced blanket replication with targeted S3 archival for critical logs only:
Critical logs archived to S3 with Intelligent-Tiering automatically move to infrequent access and then Glacier based on access patterns, costing pennies compared to live CloudWatch storage.
Step 5: Automated Log Group Governance
To prevent cost regression, we implemented automated governance that enforces our tiering policy on new log groups:
import boto3
from datetime import datetime
def enforce_log_governance(event: dict, context) -> dict:
"""EventBridge rule triggered when new log groups are created.
Applies retention, subscription filters, and alerts on non-compliant groups.
"""
logs = boto3.client('logs')
log_group_name = event['detail']['requestParameters']['logGroupName']
# Determine tier from naming convention
tier = classify_log_group(log_group_name)
# Apply retention policy
retention_days = {'critical': 365, 'standard': 30, 'debug': 3, 'ephemeral': 1}
logs.put_retention_policy(
logGroupName=log_group_name,
retentionInDays=retention_days.get(tier, 30),
)
# Tag for cost allocation
logs.tag_log_group(
logGroupName=log_group_name,
tags={
'log-tier': tier,
'cost-center': extract_team(log_group_name),
'created-date': datetime.utcnow().isoformat(),
'governance-applied': 'true',
}
)
return {'log_group': log_group_name, 'tier': tier, 'retention': retention_days[tier]}
Results: Cost Breakdown
| Category | Before | After | Savings |
|---|---|---|---|
| Ingestion volume | 820 GB/mo | 234 GB/mo | -71% |
| Stored volume | 24.6 TB | 4.1 TB | -83% |
| Cross-region replication | $4,200/mo | $0 | -100% |
| CloudWatch total | $12,400/mo | $3,200/mo | -73% |
| S3 archival (new cost) | $0 | $180/mo | New |
| Net monthly cost | $12,400 | $3,380 | -73% |
Debugging Capability Assessment
The critical question: did we lose visibility? We measured debugging effectiveness before and after:
- Incidents where logs were insufficient: Before: 4/month. After: 3/month (slight improvement from structured logging).
- Time to find relevant log entries: Before: 8.2 minutes average. After: 3.1 minutes (structured logs + CloudWatch Insights).
- Cross-service correlation: Before: manual timestamp matching. After: trace ID correlation across all logs.
Structured logging with trace IDs actually improved debugging despite ingesting 71% less data.
Lessons Learned
Health check logs are pure waste. They represent 30%+ of volume in most Kubernetes environments and have never once been useful for debugging.
Retention diversity is the highest-leverage change. Moving from uniform 30-day retention to tiered retention (1/3/30/365 days) cut storage costs by 83% with zero impact on daily operations.
Sampling is fine for dashboards. A 10% sample of successful requests provides statistically identical trend data while cutting volume 90%. Keep 100% of errors.
Team cost visibility drives behavior change. After we added per-team log cost dashboards, three teams voluntarily reduced their logging verbosity without any mandate.
Conclusion
Log cost optimization is not about logging less. It is about logging intelligently: keep everything that aids debugging, sample what is only useful for trends, and eliminate what serves no purpose. The 73% cost reduction came from applying different strategies to different log tiers, not from a single silver bullet. Start with health check elimination and retention tiering. Those two changes alone will likely save 50% or more.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.