Cutting Our Kubernetes Bill: A Practical Playbook

Concrete, ordered steps to reduce Kubernetes infrastructure costs: right-sizing, spot instances, autoscaling, and the observability to keep it that way.

#kubernetes#cost-optimization#devops#finops
Cover image for the article: Cutting Our Kubernetes Bill: A Practical Playbook

Kubernetes makes it easy to scale up, and just as easy to quietly burn money. Most clusters I've audited run at 20–35% actual utilization while paying for 100%. At Rafeeq, we cut our Kubernetes spend by 47% — from $38K/month to $20K/month — without reducing capacity or degrading performance. Here's the playbook I use, in the order that delivers results fastest.

Step 1: Measure before you touch anything

You need per-workload visibility first. Without it, you're optimizing blind. Install an open-source cost tool (OpenCost or Kubecost) or use your cloud's native tooling, and answer:

  • Which namespaces/teams spend the most?
  • What's the gap between requested and actually used CPU/memory?
  • Which workloads are idle during off-hours?
  • What percentage of your nodes are spot-eligible?

That requests-vs-usage gap is almost always your biggest win. In our case, the average workload was requesting 4x what it actually used at P95.

Step 2: Right-size requests

Engineers set resource requests defensively ("better safe than OOMKilled"), then never revisit them. A service asking for 2 CPUs while using 200m wastes 90% of its reservation, and the scheduler packs nodes based on requests, not usage. This means you're paying for nodes that are "full" according to the scheduler but mostly idle in reality.

Our right-sizing process:

  • Pull actual P95 usage from Prometheus/CloudWatch for every deployment over 7 days.
  • Set requests to 1.2x P95 usage (small buffer for spikes), limits to 2x requests.
  • Automate it: a Vertical Pod Autoscaler in recommendation mode gives you the numbers even if you apply them manually.
  • Review quarterly — traffic patterns change, and requests drift.

Result for us: This step alone cut 34% off the bill. The largest single fix was our notification service: requesting 4 CPU / 8GB RAM, actually using 0.3 CPU / 512MB at P99. Multiplied across 12 replicas, that's 44 wasted CPU cores worth of node capacity.

ServiceBefore (requests)After (requests)Savings
Notification4 CPU / 8GB0.5 CPU / 768MB87%
Order API2 CPU / 4GB1 CPU / 2GB50%
Analytics worker1 CPU / 2GB0.25 CPU / 512MB75%

Step 3: Spot/preemptible instances for everything stateless

Spot capacity is 60–90% cheaper. The rules of engagement:

  • Stateless, replicated services (3+ replicas): spot by default.
  • Batch jobs, CI runners, data processing: absolutely spot.
  • Databases, queues, anything with local state: on-demand only.
  • Anything with a single replica: on-demand (interruption = downtime).

Mix node pools and use taints/tolerations so critical pods never land on spot nodes. Modern autoscalers (Karpenter on AWS, GKE Autopilot's flexible provisioning) handle interruptions gracefully enough that this is now boring technology.

Our approach: We run 3 node pools:

  1. system — on-demand, for control plane components and stateful workloads
  2. general-spot — spot instances, for all stateless services
  3. batch-spot — spot with aggressive scaling, for background jobs

70% of our compute now runs on spot, saving approximately $11K/month.

Step 4: Autoscale in both directions

Most teams configure scale-up but forget scale-down. Your cluster should breathe:

  • Horizontal Pod Autoscaling on real signals (requests per second, queue depth), not just CPU. CPU-based HPA is a trailing indicator — by the time CPU spikes, users are already waiting.
  • Cluster autoscaler / Karpenter to shed empty nodes fast. Check your scale-down settings; the defaults are conservative and leave nodes idling for 10+ minutes.
  • Scale to zero for dev/staging environments outside working hours. A cron job that scales staging down at 8 PM and up at 8 AM pays for someone's salary in a mid-sized org. We save $4K/month just from turning off non-production clusters at night.
  • Pod Disruption Budgets (PDBs) so the autoscaler can safely drain nodes without impacting availability.

Step 5: Commit on what's left

Only after right-sizing: buy Savings Plans / Committed Use Discounts for your now-honest baseline. Committing before optimizing locks in your waste.

Our commit strategy:

  • Compute the stable baseline (minimum nodes that never scale down, even at 3 AM).
  • Buy 1-year Savings Plans for 80% of that baseline (leave 20% uncommitted for flexibility).
  • Re-evaluate quarterly as traffic patterns shift.

This saved an additional 22% on the remaining on-demand portion.

Step 6: Eliminate waste at the workload level

Beyond infrastructure-level optimization:

  • Delete unused namespaces. We found 3 staging environments from abandoned projects still running 24/7.
  • Set resource quotas per namespace. Teams can't accidentally spin up 50 replicas without hitting a guard rail.
  • Limit container image sizes. A 2GB image that could be 200MB means slower scaling and more network transfer costs.
  • Review CronJobs. That hourly job that runs for 2 seconds but provisions a full node for 10 minutes? Run it on a shared batch pool or switch to a lightweight runtime.

Keeping it optimized

Cost regressions are like performance regressions — they creep back without guardrails:

  1. Cost allocation labels enforced by admission policy. No label = no deploy. We use OPA/Gatekeeper to enforce team, environment, and service labels on every pod.
  2. A monthly cost review per team, with the requests-vs-usage gap on one dashboard. We share this in our monthly engineering all-hands.
  3. Budget alerts that page the owning team, not the platform team. When the order-service team's spend jumps 30% overnight, they investigate first.
  4. PR-level cost estimation. We're experimenting with Infracost-style checks that estimate the cost impact of Helm chart changes before merge.

The results

OptimizationMonthly savingsEffort
Right-sizing requests$12,9002 days
Spot instances (70% fleet)$11,2001 day
Night/weekend shutdown (staging)$4,1002 hours
Savings Plans (baseline)$3,8001 hour
Cleanup (dead namespaces, oversized images)$2,400Half day
Total$34,400/month~4 days

FinOps isn't a project, it's a practice. The goal isn't the lowest possible bill; it's knowing that every dollar you spend is one you chose to spend.

Comments

    No comments yet. Be the first to share your thoughts.