Reducing AWS OpenSearch Costs by 65% with Tiered Storage
How we cut OpenSearch logging costs from $18K to $6.3K/month using UltraWarm, cold storage, index lifecycle policies, and query optimization.

OpenSearch is expensive. We were spending $18,000/month on a logging cluster ingesting 2TB/day from 400 microservices. Most of that data was queried once in the first 24 hours, occasionally in the first week, and almost never after 30 days — yet we stored it all on hot SSD instances at premium rates.
By implementing tiered storage, index lifecycle policies, and selective ingestion, we reduced costs to $6,300/month — a 65% reduction — while maintaining query performance for the data that actually gets accessed.
The Problem: Flat Storage Architecture
Our OpenSearch cluster treated all logs equally: hot SSD storage with full indexing, regardless of whether the data was accessed once or a thousand times.
| Metric | Before Optimization |
|---|---|
| Daily ingestion | 2TB |
| Retention period | 90 days |
| Total stored | ~180TB |
| Instance type | r6g.2xlarge.search × 12 |
| Storage (EBS gp3) | 180TB × $0.08/GB |
| Monthly cost | $18,200 |
| Query frequency (0-24h) | 95% of queries |
| Query frequency (1-7d) | 4% of queries |
| Query frequency (7-90d) | 1% of queries |
95% of queries hit data less than 24 hours old. We were paying premium storage rates for 90 days of data that was essentially write-once-read-never.
Architecture: Three-Tier Storage Strategy
The optimized architecture uses three storage tiers aligned with access patterns:
| Tier | Age | Storage Type | Query Latency | Cost/GB/Month |
|---|---|---|---|---|
| Hot | 0-24h | SSD (instance store) | <100ms | $0.135 |
| UltraWarm | 1-30d | S3-backed managed | 1-5s | $0.024 |
| Cold | 30-90d | S3 (detached) | 30-60s | $0.011 |
// Index State Management (ISM) policy
const ismPolicy = {
policy: {
description: 'Tiered storage lifecycle for application logs',
default_state: 'hot',
states: [
{
name: 'hot',
actions: [
{
rollover: {
min_size: '50gb',
min_index_age: '1d',
},
},
],
transitions: [
{
state_name: 'warm',
conditions: { min_index_age: '1d' },
},
],
},
{
name: 'warm',
actions: [
{
warm_migration: {},
},
{
replica_count: { number_of_replicas: 0 },
},
{
force_merge: { max_num_segments: 1 },
},
],
transitions: [
{
state_name: 'cold',
conditions: { min_index_age: '30d' },
},
],
},
{
name: 'cold',
actions: [
{
cold_migration: {
timestamp_field: '@timestamp',
},
},
],
transitions: [
{
state_name: 'delete',
conditions: { min_index_age: '90d' },
},
],
},
{
name: 'delete',
actions: [{ cold_delete: {} }],
},
],
ism_template: [
{
index_patterns: ['logs-*'],
priority: 100,
},
],
},
};
Cluster Configuration (Terraform)
resource "aws_opensearch_domain" "logs" {
domain_name = "prod-logs"
engine_version = "OpenSearch_2.11"
cluster_config {
instance_type = "r6g.xlarge.search"
instance_count = 6
dedicated_master_enabled = true
dedicated_master_type = "m6g.large.search"
dedicated_master_count = 3
zone_awareness_enabled = true
zone_awareness_config {
availability_zone_count = 3
}
# UltraWarm tier
warm_enabled = true
warm_type = "ultrawarm1.large.search"
warm_count = 4
# Cold storage
cold_storage_options {
enabled = true
}
}
ebs_options {
ebs_enabled = true
volume_type = "gp3"
volume_size = 500 # Per hot node — reduced from 5TB each
throughput = 250
iops = 3000
}
advanced_options = {
"rest.action.multi.allow_explicit_index" = "true"
"indices.query.bool.max_clause_count" = "1024"
}
}
Optimization 2: Selective Ingestion
Not all log lines deserve indexing. We categorized log sources by value and implemented selective ingestion at the Fluent Bit layer:
# fluent-bit.conf - Selective ingestion pipeline
[FILTER]
Name grep
Match app.*
Exclude log ^(DEBUG|TRACE)
[FILTER]
Name modify
Match app.*
# Drop high-cardinality fields that bloat index size
Remove request_body
Remove response_body
Remove stack_trace_full
[FILTER]
Name lua
Match app.*
script /fluent-bit/scripts/sampling.lua
call sample_by_status
-- sampling.lua: Sample 200-level logs at 10%, keep all errors
function sample_by_status(tag, timestamp, record)
local status = record["http_status"] or 200
if status >= 400 then
return 0, 0, 0 -- Keep all errors
elseif status >= 200 and status < 300 then
if math.random(100) <= 10 then
record["_sampled"] = true
record["_sample_rate"] = 10
return 1, timestamp, record -- Keep 10% of success logs
else
return -1, 0, 0 -- Drop
end
end
return 0, 0, 0 -- Keep everything else
end
This reduced ingestion volume by 40% with zero impact on debugging capability — success logs are sampled, errors are always captured.
Optimization 3: Index Template Tuning
Default index settings are generous. We tuned templates for cost efficiency:
{
"index_patterns": ["logs-*"],
"template": {
"settings": {
"number_of_shards": 3,
"number_of_replicas": 1,
"codec": "zstd_no_dict",
"refresh_interval": "30s",
"translog.durability": "async",
"translog.sync_interval": "30s",
"merge.policy.max_merged_segment": "5gb"
},
"mappings": {
"dynamic": "false",
"properties": {
"@timestamp": { "type": "date" },
"level": { "type": "keyword" },
"service": { "type": "keyword" },
"message": { "type": "text", "index": true, "norms": false },
"trace_id": { "type": "keyword" },
"http_status": { "type": "short" },
"duration_ms": { "type": "float" },
"user_id": { "type": "keyword" },
"error_type": { "type": "keyword" }
}
}
}
}
Key decisions:
dynamic: false— prevents auto-mapping of unexpected fields (biggest index bloat cause)codec: zstd_no_dict— 30% better compression than default LZ4refresh_interval: 30s— reduces segment creation overhead (default is 1s)norms: falseon text fields — saves 1 byte per document per field for scoring we don't use
Cost Breakdown: Before and After
| Component | Before | After | Savings |
|---|---|---|---|
| Hot instances (r6g.2xlarge × 12) | $8,640 | $2,880 (r6g.xlarge × 6) | -67% |
| EBS storage (180TB gp3) | $7,200 | $1,200 (3TB hot only) | -83% |
| UltraWarm instances | $0 | $1,440 (ultrawarm1.large × 4) | N/A |
| Cold storage (S3) | $0 | $320 (30TB) | N/A |
| Dedicated masters | $1,440 | $460 (m6g.large × 3) | -68% |
| Data transfer | $920 | $0 (VPC endpoint) | -100% |
| Total | $18,200 | $6,300 | -65% |
Query Performance Impact
The trade-off is query latency for older data. Here's what users experience:
| Query Pattern | Before (all hot) | After (tiered) | Acceptable? |
|---|---|---|---|
| Last 1 hour, specific service | 150ms | 120ms | Yes (faster — smaller hot tier) |
| Last 24 hours, full-text search | 800ms | 600ms | Yes (improved) |
| Last 7 days, error aggregation | 1.2s | 4.5s | Yes (UltraWarm, async dashboards) |
| Last 30 days, trend analysis | 3s | 12s | Yes (scheduled reports only) |
| Last 90 days, compliance audit | 5s | 45s | Yes (rare, batch operation) |
The hot tier actually got faster because it's smaller and the instances aren't competing for I/O with 90 days of stale data.
Implementation Timeline
| Week | Action | Cost Impact |
|---|---|---|
| 1 | Enable UltraWarm, configure ISM policy | Neutral (data starts migrating) |
| 2 | Implement Fluent Bit sampling | -40% ingestion volume |
| 3 | Apply index template optimizations | -30% storage per index |
| 4 | Enable cold storage, reduce hot instance count | -65% total cost |
| 5 | Tune and validate | Final optimization |
Key Takeaways
- Align storage tier to access pattern — 95% of queries hit 1% of your data. Price accordingly.
- Sample success logs aggressively — you need every error, but 10% sampling on 200s still gives statistical validity.
- Disable dynamic mapping — it's the single biggest cause of index bloat and unexpected cost growth.
- UltraWarm for 1-30 day data — 82% cheaper than hot storage with acceptable query latency for dashboards.
- Measure before cutting — our query access pattern analysis justified the 65% cost reduction without operational pushback.
OpenSearch cost optimization isn't about using less logging — it's about matching your storage economics to your access patterns. The data that matters is recent, and your infrastructure costs should reflect that reality.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.