Reducing AWS OpenSearch Costs by 65% with Tiered Storage

How we cut OpenSearch logging costs from $18K to $6.3K/month using UltraWarm, cold storage, index lifecycle policies, and query optimization.

#aws#opensearch#logging#cost-optimization
Cover image for the article: Reducing AWS OpenSearch Costs by 65% with Tiered Storage

OpenSearch is expensive. We were spending $18,000/month on a logging cluster ingesting 2TB/day from 400 microservices. Most of that data was queried once in the first 24 hours, occasionally in the first week, and almost never after 30 days — yet we stored it all on hot SSD instances at premium rates.

By implementing tiered storage, index lifecycle policies, and selective ingestion, we reduced costs to $6,300/month — a 65% reduction — while maintaining query performance for the data that actually gets accessed.

The Problem: Flat Storage Architecture

Our OpenSearch cluster treated all logs equally: hot SSD storage with full indexing, regardless of whether the data was accessed once or a thousand times.

MetricBefore Optimization
Daily ingestion2TB
Retention period90 days
Total stored~180TB
Instance typer6g.2xlarge.search × 12
Storage (EBS gp3)180TB × $0.08/GB
Monthly cost$18,200
Query frequency (0-24h)95% of queries
Query frequency (1-7d)4% of queries
Query frequency (7-90d)1% of queries

Query Access Pattern

95% of queries hit data less than 24 hours old. We were paying premium storage rates for 90 days of data that was essentially write-once-read-never.

Architecture: Three-Tier Storage Strategy

The optimized architecture uses three storage tiers aligned with access patterns:

TierAgeStorage TypeQuery LatencyCost/GB/Month
Hot0-24hSSD (instance store)<100ms$0.135
UltraWarm1-30dS3-backed managed1-5s$0.024
Cold30-90dS3 (detached)30-60s$0.011
// Index State Management (ISM) policy
const ismPolicy = {
  policy: {
    description: 'Tiered storage lifecycle for application logs',
    default_state: 'hot',
    states: [
      {
        name: 'hot',
        actions: [
          {
            rollover: {
              min_size: '50gb',
              min_index_age: '1d',
            },
          },
        ],
        transitions: [
          {
            state_name: 'warm',
            conditions: { min_index_age: '1d' },
          },
        ],
      },
      {
        name: 'warm',
        actions: [
          {
            warm_migration: {},
          },
          {
            replica_count: { number_of_replicas: 0 },
          },
          {
            force_merge: { max_num_segments: 1 },
          },
        ],
        transitions: [
          {
            state_name: 'cold',
            conditions: { min_index_age: '30d' },
          },
        ],
      },
      {
        name: 'cold',
        actions: [
          {
            cold_migration: {
              timestamp_field: '@timestamp',
            },
          },
        ],
        transitions: [
          {
            state_name: 'delete',
            conditions: { min_index_age: '90d' },
          },
        ],
      },
      {
        name: 'delete',
        actions: [{ cold_delete: {} }],
      },
    ],
    ism_template: [
      {
        index_patterns: ['logs-*'],
        priority: 100,
      },
    ],
  },
};

Cluster Configuration (Terraform)

resource "aws_opensearch_domain" "logs" {
  domain_name    = "prod-logs"
  engine_version = "OpenSearch_2.11"

  cluster_config {
    instance_type            = "r6g.xlarge.search"
    instance_count           = 6
    dedicated_master_enabled = true
    dedicated_master_type    = "m6g.large.search"
    dedicated_master_count   = 3
    zone_awareness_enabled   = true

    zone_awareness_config {
      availability_zone_count = 3
    }

    # UltraWarm tier
    warm_enabled = true
    warm_type    = "ultrawarm1.large.search"
    warm_count   = 4

    # Cold storage
    cold_storage_options {
      enabled = true
    }
  }

  ebs_options {
    ebs_enabled = true
    volume_type = "gp3"
    volume_size = 500  # Per hot node — reduced from 5TB each
    throughput  = 250
    iops        = 3000
  }

  advanced_options = {
    "rest.action.multi.allow_explicit_index" = "true"
    "indices.query.bool.max_clause_count"    = "1024"
  }
}

Optimization 2: Selective Ingestion

Not all log lines deserve indexing. We categorized log sources by value and implemented selective ingestion at the Fluent Bit layer:

# fluent-bit.conf - Selective ingestion pipeline
[FILTER]
    Name    grep
    Match   app.*
    Exclude log ^(DEBUG|TRACE)

[FILTER]
    Name    modify
    Match   app.*
    # Drop high-cardinality fields that bloat index size
    Remove  request_body
    Remove  response_body
    Remove  stack_trace_full

[FILTER]
    Name    lua
    Match   app.*
    script  /fluent-bit/scripts/sampling.lua
    call    sample_by_status
-- sampling.lua: Sample 200-level logs at 10%, keep all errors
function sample_by_status(tag, timestamp, record)
    local status = record["http_status"] or 200
    if status >= 400 then
        return 0, 0, 0  -- Keep all errors
    elseif status >= 200 and status &#x3C; 300 then
        if math.random(100) &#x3C;= 10 then
            record["_sampled"] = true
            record["_sample_rate"] = 10
            return 1, timestamp, record  -- Keep 10% of success logs
        else
            return -1, 0, 0  -- Drop
        end
    end
    return 0, 0, 0  -- Keep everything else
end

This reduced ingestion volume by 40% with zero impact on debugging capability — success logs are sampled, errors are always captured.

Ingestion Reduction Breakdown

Optimization 3: Index Template Tuning

Default index settings are generous. We tuned templates for cost efficiency:

{
  "index_patterns": ["logs-*"],
  "template": {
    "settings": {
      "number_of_shards": 3,
      "number_of_replicas": 1,
      "codec": "zstd_no_dict",
      "refresh_interval": "30s",
      "translog.durability": "async",
      "translog.sync_interval": "30s",
      "merge.policy.max_merged_segment": "5gb"
    },
    "mappings": {
      "dynamic": "false",
      "properties": {
        "@timestamp": { "type": "date" },
        "level": { "type": "keyword" },
        "service": { "type": "keyword" },
        "message": { "type": "text", "index": true, "norms": false },
        "trace_id": { "type": "keyword" },
        "http_status": { "type": "short" },
        "duration_ms": { "type": "float" },
        "user_id": { "type": "keyword" },
        "error_type": { "type": "keyword" }
      }
    }
  }
}

Key decisions:

  • dynamic: false — prevents auto-mapping of unexpected fields (biggest index bloat cause)
  • codec: zstd_no_dict — 30% better compression than default LZ4
  • refresh_interval: 30s — reduces segment creation overhead (default is 1s)
  • norms: false on text fields — saves 1 byte per document per field for scoring we don't use

Cost Breakdown: Before and After

ComponentBeforeAfterSavings
Hot instances (r6g.2xlarge × 12)$8,640$2,880 (r6g.xlarge × 6)-67%
EBS storage (180TB gp3)$7,200$1,200 (3TB hot only)-83%
UltraWarm instances$0$1,440 (ultrawarm1.large × 4)N/A
Cold storage (S3)$0$320 (30TB)N/A
Dedicated masters$1,440$460 (m6g.large × 3)-68%
Data transfer$920$0 (VPC endpoint)-100%
Total$18,200$6,300-65%

Query Performance Impact

The trade-off is query latency for older data. Here's what users experience:

Query PatternBefore (all hot)After (tiered)Acceptable?
Last 1 hour, specific service150ms120msYes (faster — smaller hot tier)
Last 24 hours, full-text search800ms600msYes (improved)
Last 7 days, error aggregation1.2s4.5sYes (UltraWarm, async dashboards)
Last 30 days, trend analysis3s12sYes (scheduled reports only)
Last 90 days, compliance audit5s45sYes (rare, batch operation)

The hot tier actually got faster because it's smaller and the instances aren't competing for I/O with 90 days of stale data.

Implementation Timeline

WeekActionCost Impact
1Enable UltraWarm, configure ISM policyNeutral (data starts migrating)
2Implement Fluent Bit sampling-40% ingestion volume
3Apply index template optimizations-30% storage per index
4Enable cold storage, reduce hot instance count-65% total cost
5Tune and validateFinal optimization

Key Takeaways

  1. Align storage tier to access pattern — 95% of queries hit 1% of your data. Price accordingly.
  2. Sample success logs aggressively — you need every error, but 10% sampling on 200s still gives statistical validity.
  3. Disable dynamic mapping — it's the single biggest cause of index bloat and unexpected cost growth.
  4. UltraWarm for 1-30 day data — 82% cheaper than hot storage with acceptable query latency for dashboards.
  5. Measure before cutting — our query access pattern analysis justified the 65% cost reduction without operational pushback.

OpenSearch cost optimization isn't about using less logging — it's about matching your storage economics to your access patterns. The data that matters is recent, and your infrastructure costs should reflect that reality.

Comments

    No comments yet. Be the first to share your thoughts.