Reducing CloudWatch Logs Costs from $12K to $3.2K/Month Without Losing Visibility

A systematic approach to log aggregation cost optimization using tiered retention, sampling, and structured logging that cut our bill by 73%

#logging#cost-optimization#observability#aws
Cover image for the article: Reducing CloudWatch Logs Costs from $12K to $3.2K/Month Without Losing Visibility

Our CloudWatch Logs bill hit $12,400/month and was growing 15% month-over-month. The root cause was not excessive infrastructure but undisciplined logging: verbose debug logs in production, duplicated log streams, and uniform 30-day retention across all log groups regardless of value. We implemented a tiered logging strategy that reduced costs to $3,200/month while actually improving our ability to diagnose production issues.

The Problem: All Logs Treated Equal

Every log line has a cost. At $0.50 per GB ingested and $0.03 per GB stored, CloudWatch Logs becomes expensive quickly when you are ingesting 820 GB per month without discrimination.

Cost breakdown before optimization:

  • Ingestion (820 GB x $0.50): $410/month
  • Storage (24.6 TB x $0.03): $738/month
  • But wait: we had cross-region replication enabled on all log groups
  • Total with replication and data transfer: $12,400/month

The problem was cultural as much as technical. Engineers added console.log liberally during development and never removed it. Health check endpoints logged every request. HTTP response bodies were logged at INFO level. Nobody owned log hygiene.

Architecture: Tiered Log Strategy

We classified all logs into four tiers based on their debugging value and compliance requirements:

Log Tiering Architecture

TierDescriptionRetentionStorageExample
CriticalSecurity, compliance, errors365 daysCloudWatch + S3Auth failures, 5xx errors
StandardApplication business logic30 daysCloudWatchRequest logs, state transitions
DebugDevelopment diagnostics3 daysCloudWatchVerbose debugging, SQL queries
EphemeralHealth checks, heartbeats1 dayFiltered out pre-ingestion/health, /ready, keep-alives

Step 1: Structured Logging Standard

The foundation of cost optimization is structured logging. Unstructured text logs cannot be filtered efficiently. We deployed a shared logging library across all services:

import pino from 'pino';

interface LogConfig {
  service: string;
  environment: string;
  tier?: 'critical' | 'standard' | 'debug' | 'ephemeral';
}

export function createLogger(config: LogConfig) {
  return pino({
    level: process.env.LOG_LEVEL || 'info',
    formatters: {
      level(label: string) {
        return { level: label };
      },
    },
    base: {
      service: config.service,
      environment: config.environment,
      tier: config.tier || 'standard',
      version: process.env.APP_VERSION,
    },
    // Redact sensitive fields automatically
    redact: {
      paths: [
        'req.headers.authorization',
        'req.headers.cookie',
        'body.password',
        'body.token',
        'body.creditCard',
      ],
      censor: '[REDACTED]',
    },
    // Custom serializers for cost-efficient logging
    serializers: {
      req: (req) => ({
        method: req.method,
        url: req.url,
        // Don't log full headers - saves ~200 bytes per request
        contentLength: req.headers['content-length'],
        userAgent: req.headers['user-agent'],
      }),
      res: (res) => ({
        statusCode: res.statusCode,
        // Don't log response bodies at info level
        contentLength: res.headers?.['content-length'],
      }),
    },
  });
}

// Usage
const logger = createLogger({ service: 'order-service', environment: 'production' });

// Tier-aware logging
logger.info({ tier: 'standard', orderId, action: 'created' }, 'Order created');
logger.error({ tier: 'critical', orderId, error: err.message }, 'Payment failed');
logger.debug({ tier: 'debug', query, params, duration }, 'Database query executed');

Structured logging with tier annotations enabled automated routing: critical logs go to long-term storage, debug logs get short retention, and ephemeral logs get dropped before ingestion.

Step 2: Pre-Ingestion Filtering with Subscription Filters

The biggest cost win was preventing low-value logs from ever reaching CloudWatch. We use CloudWatch subscription filters to route logs before they incur ingestion charges:

import boto3

def configure_log_group_filters(log_group_name: str, service_config: dict):
    """Configure subscription filters for cost-optimized log routing."""
    logs = boto3.client('logs')
    
    # Filter out health check logs entirely (saves ~30% of volume)
    logs.put_subscription_filter(
        logGroupName=log_group_name,
        filterName='drop-health-checks',
        filterPattern='{ $.url != "/health" && $.url != "/ready" && $.url != "/metrics" }',
        destinationArn=service_config['firehose_arn'],
        roleArn=service_config['role_arn'],
    )
    
    # Set retention based on tier
    retention_map = {
        'critical': 365,
        'standard': 30,
        'debug': 3,
        'ephemeral': 1,
    }
    
    tier = service_config.get('tier', 'standard')
    logs.put_retention_policy(
        logGroupName=log_group_name,
        retentionInDays=retention_map[tier],
    )
    
    # Enable S3 archival for critical logs
    if tier == 'critical':
        logs.put_subscription_filter(
            logGroupName=log_group_name,
            filterName='archive-to-s3',
            filterPattern='{ $.tier = "critical" }',
            destinationArn=service_config['s3_firehose_arn'],
            roleArn=service_config['role_arn'],
        )

Health check filtering alone saved 30% of our ingestion volume. These endpoints are called every 10 seconds by load balancers and Kubernetes probes, generating massive log volume with zero debugging value.

Step 3: Log Sampling for High-Volume Endpoints

Some endpoints generate legitimate logs at volumes that are cost-prohibitive to retain at 100%. We implemented probabilistic sampling for successful requests while keeping 100% of errors:

import { createLogger } from './logger';

const logger = createLogger({ service: 'api-gateway', environment: 'production' });

function shouldSample(statusCode: number, path: string): boolean {
  // Always log errors
  if (statusCode >= 400) return true;
  
  // Always log slow requests
  // (checked elsewhere via response time)
  
  // Sample high-volume endpoints at 10%
  const highVolumeEndpoints = ['/api/v1/products', '/api/v1/search', '/api/v1/recommendations'];
  if (highVolumeEndpoints.some(ep => path.startsWith(ep))) {
    return Math.random() < 0.1;
  }
  
  // Sample all other successful requests at 25%
  return Math.random() < 0.25;
}

export function requestLogger(req: Request, res: Response, next: () => void) {
  const start = Date.now();
  
  res.on('finish', () => {
    const duration = Date.now() - start;
    const shouldLog = shouldSample(res.statusCode, req.path) || duration > 2000;
    
    if (shouldLog) {
      logger.info({
        tier: res.statusCode >= 500 ? 'critical' : 'standard',
        method: req.method,
        path: req.path,
        statusCode: res.statusCode,
        duration,
        sampled: true,
        sampleRate: getSampleRate(req.path),
      }, 'Request completed');
    }
  });
  
  next();
}

Sampling at 10-25% for successful requests provides sufficient statistical data for dashboards and trend analysis while dramatically reducing volume.

Step 4: Cross-Region Replication Elimination

Our largest unnecessary cost was cross-region log replication. Every log group was replicated to a DR region "just in case." Analysis showed that DR log access happened zero times in 18 months. We replaced blanket replication with targeted S3 archival for critical logs only:

Log Storage Architecture

Critical logs archived to S3 with Intelligent-Tiering automatically move to infrequent access and then Glacier based on access patterns, costing pennies compared to live CloudWatch storage.

Step 5: Automated Log Group Governance

To prevent cost regression, we implemented automated governance that enforces our tiering policy on new log groups:

import boto3
from datetime import datetime

def enforce_log_governance(event: dict, context) -> dict:
    """EventBridge rule triggered when new log groups are created.
    
    Applies retention, subscription filters, and alerts on non-compliant groups.
    """
    logs = boto3.client('logs')
    log_group_name = event['detail']['requestParameters']['logGroupName']
    
    # Determine tier from naming convention
    tier = classify_log_group(log_group_name)
    
    # Apply retention policy
    retention_days = {'critical': 365, 'standard': 30, 'debug': 3, 'ephemeral': 1}
    logs.put_retention_policy(
        logGroupName=log_group_name,
        retentionInDays=retention_days.get(tier, 30),
    )
    
    # Tag for cost allocation
    logs.tag_log_group(
        logGroupName=log_group_name,
        tags={
            'log-tier': tier,
            'cost-center': extract_team(log_group_name),
            'created-date': datetime.utcnow().isoformat(),
            'governance-applied': 'true',
        }
    )
    
    return {'log_group': log_group_name, 'tier': tier, 'retention': retention_days[tier]}

Results: Cost Breakdown

CategoryBeforeAfterSavings
Ingestion volume820 GB/mo234 GB/mo-71%
Stored volume24.6 TB4.1 TB-83%
Cross-region replication$4,200/mo$0-100%
CloudWatch total$12,400/mo$3,200/mo-73%
S3 archival (new cost)$0$180/moNew
Net monthly cost$12,400$3,380-73%

Debugging Capability Assessment

The critical question: did we lose visibility? We measured debugging effectiveness before and after:

  • Incidents where logs were insufficient: Before: 4/month. After: 3/month (slight improvement from structured logging).
  • Time to find relevant log entries: Before: 8.2 minutes average. After: 3.1 minutes (structured logs + CloudWatch Insights).
  • Cross-service correlation: Before: manual timestamp matching. After: trace ID correlation across all logs.

Structured logging with trace IDs actually improved debugging despite ingesting 71% less data.

Lessons Learned

Health check logs are pure waste. They represent 30%+ of volume in most Kubernetes environments and have never once been useful for debugging.

Retention diversity is the highest-leverage change. Moving from uniform 30-day retention to tiered retention (1/3/30/365 days) cut storage costs by 83% with zero impact on daily operations.

Sampling is fine for dashboards. A 10% sample of successful requests provides statistically identical trend data while cutting volume 90%. Keep 100% of errors.

Team cost visibility drives behavior change. After we added per-team log cost dashboards, three teams voluntarily reduced their logging verbosity without any mandate.

Conclusion

Log cost optimization is not about logging less. It is about logging intelligently: keep everything that aids debugging, sample what is only useful for trends, and eliminate what serves no purpose. The 73% cost reduction came from applying different strategies to different log tiers, not from a single silver bullet. Start with health check elimination and retention tiering. Those two changes alone will likely save 50% or more.

Comments

    No comments yet. Be the first to share your thoughts.