GCP Cloud Logging Cost Control and Optimization

Strategies to reduce Google Cloud Logging costs by 50-80% through exclusion filters, log routing, retention tuning, and sampling without losing critical observability

#gcp#cloud-logging#cost-optimization#observability
Cover image for the article: GCP Cloud Logging Cost Control and Optimization

Introduction

Google Cloud Logging ingestion costs are one of the most surprising line items in GCP bills. At $0.50 per GB ingested, a cluster generating 100 GB of logs daily costs $1,500 per month just for log ingestion. Most organizations log far more than they need, with studies showing that 60-80% of ingested logs are never queried or used for alerting.

This article provides a systematic approach to reducing Cloud Logging costs while maintaining the observability necessary for debugging, auditing, and compliance.

Understanding the Cost Structure

ComponentPriceFree TierNotes
Ingestion$0.50/GB50 GB/project/monthAll logs routed to _Default bucket
Storage (> 30 days)$0.01/GB/month30 days retention freePer log bucket
Log RouterFreeN/ANo charge for routing decisions
Exclusion filtersFreeN/AReduces ingestion costs
Log Analytics$0.01/GB scannedN/ABigQuery-compatible queries

Chart

Where Costs Accumulate

Log SourceTypical Volume (per 100 nodes)Monthly Cost% of Total
GKE container stdout/stderr60 GB/day$900/mo45%
Load balancer access logs30 GB/day$450/mo22%
VPC flow logs25 GB/day$375/mo18%
Cloud Audit logs (data access)15 GB/day$225/mo11%
Application logs5 GB/day$75/mo4%
Total135 GB/day$2,025/mo100%

Strategy 1: Exclusion Filters

Exclusion filters drop logs before ingestion, meaning you pay nothing for excluded logs:

# Exclude GKE health check logs (high volume, zero value)
gcloud logging sinks update _Default \
  --log-filter='NOT (
    resource.type="k8s_container" AND
    httpRequest.requestUrl="/healthz" OR
    httpRequest.requestUrl="/readyz" OR
    httpRequest.requestUrl="/livez"
  )'

# Create exclusion filter for verbose debug logs
gcloud logging exclusions create "exclude-debug-logs" \
  --description="Exclude DEBUG level container logs" \
  --log-filter='resource.type="k8s_container" AND severity="DEBUG"'

# Exclude high-volume load balancer 2xx logs (keep errors)
gcloud logging exclusions create "exclude-lb-success" \
  --description="Exclude successful LB requests" \
  --log-filter='resource.type="http_load_balancer" AND httpRequest.status>=200 AND httpRequest.status<300'

# Exclude noisy system logs
gcloud logging exclusions create "exclude-gke-system" \
  --description="Exclude kube-system namespace verbose logs" \
  --log-filter='resource.type="k8s_container" AND resource.labels.namespace_name="kube-system" AND severity<="INFO"'

Exclusion Filter Impact

FilterVolume ExcludedMonthly SavingsRisk Level
Health check logs15 GB/day$225/moNone
DEBUG-level logs12 GB/day$180/moLow
LB 2xx access logs20 GB/day$300/moLow-Medium
kube-system INFO8 GB/day$120/moLow
VPC flow logs (sampled)20 GB/day$300/moMedium
Total75 GB/day$1,125/mo

Strategy 2: Log Routing to Cheaper Destinations

Route logs to BigQuery or Cloud Storage instead of Cloud Logging for long-term analysis:

# Route to BigQuery for analytics (cheaper long-term storage)
gcloud logging sinks create "bq-audit-logs" \
  bigquery.googleapis.com/projects/my-project/datasets/audit_logs \
  --log-filter='protoPayload.@type="type.googleapis.com/google.cloud.audit.v1.AuditLog"' \
  --description="Route audit logs to BigQuery for long-term analysis"

# Route to Cloud Storage for archival
gcloud logging sinks create "gcs-archive" \
  storage.googleapis.com/my-log-archive-bucket \
  --log-filter='resource.type="k8s_container" AND severity>="WARNING"' \
  --description="Archive warning+ logs to GCS"

# Route to Pub/Sub for real-time processing
gcloud logging sinks create "pubsub-errors" \
  pubsub.googleapis.com/projects/my-project/topics/error-logs \
  --log-filter='severity>="ERROR"' \
  --description="Stream errors to alerting pipeline"

Cost Comparison by Destination

DestinationIngestionStorage (30 days, 100 GB)Query CostTotal (100 GB/mo)
Cloud Logging$50Free (< 30d)Free (Logs Explorer)$50
BigQueryFree$0.20$5/TB scanned$0.70
Cloud Storage (Standard)Free$2.00N/A (export to BQ)$2.00
Cloud Storage (Coldline)Free$0.40$2 retrieval$0.40

Strategy 3: Retention Tuning

Reduce retention from default 30 days to match actual needs:

# Create custom log bucket with reduced retention
gcloud logging buckets create "short-term-logs" \
  --location=global \
  --retention-days=7 \
  --description="7-day retention for high-volume operational logs"

# Create bucket for compliance logs with extended retention
gcloud logging buckets create "compliance-logs" \
  --location=global \
  --retention-days=365 \
  --description="1-year retention for audit and compliance"

# Update _Default bucket retention
gcloud logging buckets update _Default \
  --location=global \
  --retention-days=14
Log TypeRecommended RetentionRationale
Application debug3-7 daysOnly needed during active debugging
Access logs14 daysPerformance analysis window
Error logs30 daysIncident investigation window
Audit logs (admin)365 daysCompliance requirement
Audit logs (data access)90 daysSecurity investigation
VPC flow logs14 daysNetwork forensics

Strategy 4: Log Sampling

For high-volume logs where statistical representation suffices:

import google.cloud.logging
import random

client = google.cloud.logging.Client()
logger = client.logger("my-application")

SAMPLE_RATE = 0.1  # Log 10% of requests

def log_request(request_data):
    """Sample non-error request logs at 10%."""
    if request_data.get('status_code', 200) >= 400:
        # Always log errors
        logger.log_struct(request_data, severity='ERROR')
    elif random.random() &#x3C; SAMPLE_RATE:
        # Sample successful requests
        request_data['_sampled'] = True
        request_data['_sample_rate'] = SAMPLE_RATE
        logger.log_struct(request_data, severity='INFO')

VPC Flow Log Sampling

# Enable VPC flow logs with sampling (reduce volume by 50%)
gcloud compute networks subnets update my-subnet \
  --region=us-central1 \
  --enable-flow-logs \
  --logging-flow-sampling=0.5 \
  --logging-aggregation-interval=INTERVAL_10_MIN \
  --logging-metadata=INCLUDE_ALL_METADATA

Strategy 5: Structured Logging for Efficient Querying

Structured logs reduce query costs and improve filter precision:

import google.cloud.logging
import json

client = google.cloud.logging.Client()
logger = client.logger("order-service")

def log_order_event(order_id, event_type, details):
    """Structured log entry for efficient filtering."""
    logger.log_struct({
        "order_id": order_id,
        "event_type": event_type,
        "details": details,
        "service": "order-service",
        "version": "2.1.0"
    }, severity='INFO')

# This enables precise exclusion filters like:
# resource.type="k8s_container" AND
# jsonPayload.event_type!="health_check" AND
# jsonPayload.service="order-service"

Complete Optimization Implementation

# Terraform: Complete logging optimization

# Custom log buckets with appropriate retention
resource "google_logging_project_bucket_config" "operational" {
  project        = var.project_id
  location       = "global"
  retention_days = 7
  bucket_id      = "operational-logs"
}

resource "google_logging_project_bucket_config" "security" {
  project        = var.project_id
  location       = "global"
  retention_days = 365
  bucket_id      = "security-logs"
}

# Exclusion filters
resource "google_logging_project_exclusion" "health_checks" {
  name        = "exclude-health-checks"
  description = "Exclude health check request logs"
  filter      = &#x3C;&#x3C;-EOT
    resource.type="k8s_container" AND
    (httpRequest.requestUrl="/healthz" OR
     httpRequest.requestUrl="/readyz" OR
     httpRequest.requestUrl="/livez")
  EOT
}

resource "google_logging_project_exclusion" "debug_logs" {
  name        = "exclude-debug-logs"
  description = "Exclude DEBUG severity logs"
  filter      = "severity=\"DEBUG\""
}

# Route security logs to dedicated bucket
resource "google_logging_project_sink" "security_sink" {
  name                   = "security-log-sink"
  destination            = "logging.googleapis.com/projects/${var.project_id}/locations/global/buckets/security-logs"
  filter                 = "protoPayload.@type=\"type.googleapis.com/google.cloud.audit.v1.AuditLog\""
  unique_writer_identity = true
}

Total Savings Summary

StrategyBeforeAfterMonthly Savings% Reduction
Exclusion filters$2,025$900$1,12556%
Log routing (BQ/GCS)$900$650$25028%
Retention tuning$650$580$7011%
Sampling$580$480$10017%
Combined$2,025$480$1,54576%

Key Takeaways

  • Exclusion filters are the highest-impact optimization providing 50-60% cost reduction by dropping health checks, debug logs, and successful LB requests before ingestion.
  • Route compliance and audit logs to BigQuery where long-term storage costs $0.20/100GB instead of $1/100GB in Cloud Logging with better query capabilities.
  • Reduce _Default bucket retention to 14 days and create dedicated buckets with appropriate retention for different log categories.
  • Sample high-volume logs at 10-50% for access logs and VPC flow logs where statistical representation is sufficient for analysis.
  • Always log errors at 100% regardless of sampling configuration to ensure complete incident visibility.
  • Structured logging enables precise exclusion because you can filter on specific JSON fields rather than using regex on unstructured text.
  • Expected savings of 50-80% are achievable without losing critical observability by combining exclusion, routing, retention, and sampling strategies.

Comments

    No comments yet. Be the first to share your thoughts.