Cutting Our Kubernetes Bill: A Practical Playbook
Concrete, ordered steps to reduce Kubernetes infrastructure costs: right-sizing, spot instances, autoscaling, and the observability to keep it that way.

Kubernetes makes it easy to scale up, and just as easy to quietly burn money. Most clusters I've audited run at 20–35% actual utilization while paying for 100%. At Rafeeq, we cut our Kubernetes spend by 47% — from $38K/month to $20K/month — without reducing capacity or degrading performance. Here's the playbook I use, in the order that delivers results fastest.
Step 1: Measure before you touch anything
You need per-workload visibility first. Without it, you're optimizing blind. Install an open-source cost tool (OpenCost or Kubecost) or use your cloud's native tooling, and answer:
- Which namespaces/teams spend the most?
- What's the gap between requested and actually used CPU/memory?
- Which workloads are idle during off-hours?
- What percentage of your nodes are spot-eligible?
That requests-vs-usage gap is almost always your biggest win. In our case, the average workload was requesting 4x what it actually used at P95.
Step 2: Right-size requests
Engineers set resource requests defensively ("better safe than OOMKilled"), then never revisit them. A service asking for 2 CPUs while using 200m wastes 90% of its reservation, and the scheduler packs nodes based on requests, not usage. This means you're paying for nodes that are "full" according to the scheduler but mostly idle in reality.
Our right-sizing process:
- Pull actual P95 usage from Prometheus/CloudWatch for every deployment over 7 days.
- Set requests to 1.2x P95 usage (small buffer for spikes), limits to 2x requests.
- Automate it: a Vertical Pod Autoscaler in recommendation mode gives you the numbers even if you apply them manually.
- Review quarterly — traffic patterns change, and requests drift.
Result for us: This step alone cut 34% off the bill. The largest single fix was our notification service: requesting 4 CPU / 8GB RAM, actually using 0.3 CPU / 512MB at P99. Multiplied across 12 replicas, that's 44 wasted CPU cores worth of node capacity.
| Service | Before (requests) | After (requests) | Savings |
|---|---|---|---|
| Notification | 4 CPU / 8GB | 0.5 CPU / 768MB | 87% |
| Order API | 2 CPU / 4GB | 1 CPU / 2GB | 50% |
| Analytics worker | 1 CPU / 2GB | 0.25 CPU / 512MB | 75% |
Step 3: Spot/preemptible instances for everything stateless
Spot capacity is 60–90% cheaper. The rules of engagement:
- Stateless, replicated services (3+ replicas): spot by default.
- Batch jobs, CI runners, data processing: absolutely spot.
- Databases, queues, anything with local state: on-demand only.
- Anything with a single replica: on-demand (interruption = downtime).
Mix node pools and use taints/tolerations so critical pods never land on spot nodes. Modern autoscalers (Karpenter on AWS, GKE Autopilot's flexible provisioning) handle interruptions gracefully enough that this is now boring technology.
Our approach: We run 3 node pools:
system— on-demand, for control plane components and stateful workloadsgeneral-spot— spot instances, for all stateless servicesbatch-spot— spot with aggressive scaling, for background jobs
70% of our compute now runs on spot, saving approximately $11K/month.
Step 4: Autoscale in both directions
Most teams configure scale-up but forget scale-down. Your cluster should breathe:
- Horizontal Pod Autoscaling on real signals (requests per second, queue depth), not just CPU. CPU-based HPA is a trailing indicator — by the time CPU spikes, users are already waiting.
- Cluster autoscaler / Karpenter to shed empty nodes fast. Check your scale-down settings; the defaults are conservative and leave nodes idling for 10+ minutes.
- Scale to zero for dev/staging environments outside working hours. A cron job that scales staging down at 8 PM and up at 8 AM pays for someone's salary in a mid-sized org. We save $4K/month just from turning off non-production clusters at night.
- Pod Disruption Budgets (PDBs) so the autoscaler can safely drain nodes without impacting availability.
Step 5: Commit on what's left
Only after right-sizing: buy Savings Plans / Committed Use Discounts for your now-honest baseline. Committing before optimizing locks in your waste.
Our commit strategy:
- Compute the stable baseline (minimum nodes that never scale down, even at 3 AM).
- Buy 1-year Savings Plans for 80% of that baseline (leave 20% uncommitted for flexibility).
- Re-evaluate quarterly as traffic patterns shift.
This saved an additional 22% on the remaining on-demand portion.
Step 6: Eliminate waste at the workload level
Beyond infrastructure-level optimization:
- Delete unused namespaces. We found 3 staging environments from abandoned projects still running 24/7.
- Set resource quotas per namespace. Teams can't accidentally spin up 50 replicas without hitting a guard rail.
- Limit container image sizes. A 2GB image that could be 200MB means slower scaling and more network transfer costs.
- Review CronJobs. That hourly job that runs for 2 seconds but provisions a full node for 10 minutes? Run it on a shared batch pool or switch to a lightweight runtime.
Keeping it optimized
Cost regressions are like performance regressions — they creep back without guardrails:
- Cost allocation labels enforced by admission policy. No label = no deploy. We use OPA/Gatekeeper to enforce team, environment, and service labels on every pod.
- A monthly cost review per team, with the requests-vs-usage gap on one dashboard. We share this in our monthly engineering all-hands.
- Budget alerts that page the owning team, not the platform team. When the order-service team's spend jumps 30% overnight, they investigate first.
- PR-level cost estimation. We're experimenting with Infracost-style checks that estimate the cost impact of Helm chart changes before merge.
The results
| Optimization | Monthly savings | Effort |
|---|---|---|
| Right-sizing requests | $12,900 | 2 days |
| Spot instances (70% fleet) | $11,200 | 1 day |
| Night/weekend shutdown (staging) | $4,100 | 2 hours |
| Savings Plans (baseline) | $3,800 | 1 hour |
| Cleanup (dead namespaces, oversized images) | $2,400 | Half day |
| Total | $34,400/month | ~4 days |
FinOps isn't a project, it's a practice. The goal isn't the lowest possible bill; it's knowing that every dollar you spend is one you chose to spend.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.