Tracking Cost-Per-Request Across Every Microservice
How to implement unit economics tracking for cloud infrastructure, allocating real costs to individual services and API endpoints to drive optimization decisions.

The Problem: Cloud Costs Without Attribution
Our monthly AWS bill grew from $47K to $183K in 18 months. Everyone agreed costs were too high, but nobody could answer the basic question: which services cost the most per request, and why?
Without unit economics, cost optimization is guesswork. Teams optimize what's visible (that one big EC2 instance) while ignoring systemic inefficiencies (a service making 47 redundant database calls per request). We needed a system that attributes every dollar of cloud spend to the requests it serves.
After implementing cost-per-request tracking, we identified $41K/month in optimization opportunities within the first 30 days. Six months later, our cost-per-million-requests dropped 62% while traffic grew 40%.
The Unit Economics Framework
The core metric: cost per 1 million requests (CPM) for each service. This normalizes cost against traffic and makes services comparable regardless of scale.
CPM = (Monthly Infrastructure Cost for Service) / (Monthly Requests / 1,000,000)
But "infrastructure cost for service" isn't straightforward in shared environments. We break it into three allocation layers:
- Direct costs: Resources exclusively used by one service (dedicated pods, databases)
- Shared costs: Resources used by multiple services (load balancers, shared caches, ingress)
- Platform costs: Cluster overhead (control plane, monitoring, networking)
Implementation: Cost Allocation Pipeline
Step 1: Tag Everything
Every resource must be attributed to a service. We enforce tagging through admission webhooks:
# admission-webhook/tag-enforcement.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingWebhookConfiguration
metadata:
name: cost-allocation-tagger
webhooks:
- name: cost-tagger.finops.internal
rules:
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets"]
operations: ["CREATE", "UPDATE"]
clientConfig:
service:
name: cost-tagger
namespace: finops
---
# The webhook injects these labels if missing:
# app.kubernetes.io/name: <service-name>
# finops.internal/cost-center: <team>
# finops.internal/environment: <env>
Step 2: Collect Resource Usage Per Service
We use Kubecost's allocation API combined with Prometheus metrics:
# finops/cost_collector.py
from dataclasses import dataclass
from datetime import datetime
@dataclass
class ServiceCost:
service: str
namespace: str
period: str # e.g., "2026-08"
# Direct compute costs
cpu_cost: float # Based on actual usage, not requests
memory_cost: float
gpu_cost: float
# Storage costs
pv_cost: float
s3_cost: float
# Network costs
egress_cost: float
lb_cost: float
# Shared cost allocation
shared_infra_cost: float
platform_overhead_cost: float
# Traffic
total_requests: int
@property
def total_cost(self) -> float:
return (
self.cpu_cost + self.memory_cost + self.gpu_cost +
self.pv_cost + self.s3_cost +
self.egress_cost + self.lb_cost +
self.shared_infra_cost + self.platform_overhead_cost
)
@property
def cost_per_million_requests(self) -> float:
if self.total_requests == 0:
return 0
return (self.total_cost / self.total_requests) * 1_000_000
@property
def cost_breakdown(self) -> dict[str, float]:
"""Percentage breakdown by cost category."""
total = self.total_cost
if total == 0:
return {}
return {
"compute": (self.cpu_cost + self.memory_cost + self.gpu_cost) / total * 100,
"storage": (self.pv_cost + self.s3_cost) / total * 100,
"network": (self.egress_cost + self.lb_cost) / total * 100,
"shared": (self.shared_infra_cost + self.platform_overhead_cost) / total * 100,
}
Step 3: Shared Cost Allocation
Shared resources (ALBs, NAT gateways, monitoring infrastructure) are allocated proportionally based on usage:
# finops/shared_allocation.py
def allocate_shared_costs(
shared_resource_cost: float,
resource_type: str,
service_usage: dict[str, float],
) -> dict[str, float]:
"""
Allocate shared costs proportionally based on usage metrics.
Allocation strategies by resource type:
- load_balancer: by request count
- nat_gateway: by egress bytes
- monitoring: by time series count
- platform: by cpu_seconds consumed
"""
total_usage = sum(service_usage.values())
if total_usage == 0:
# Equal split if no usage data
equal_share = shared_resource_cost / len(service_usage)
return {svc: equal_share for svc in service_usage}
return {
service: shared_resource_cost * (usage / total_usage)
for service, usage in service_usage.items()
}
Step 4: Prometheus Metrics Export
We export CPM as a Prometheus metric for dashboarding and alerting:
# finops/metrics_exporter.py
from prometheus_client import Gauge, start_http_server
cost_per_million = Gauge(
'service_cost_per_million_requests',
'Cost in USD per 1M requests',
['service', 'namespace', 'cost_category']
)
cost_total_monthly = Gauge(
'service_cost_total_monthly_usd',
'Total monthly cost in USD',
['service', 'namespace']
)
cost_efficiency_score = Gauge(
'service_cost_efficiency_score',
'Efficiency score (0-100, higher is better)',
['service', 'namespace']
)
def update_metrics(services: list[ServiceCost]):
for svc in services:
cost_per_million.labels(
service=svc.service,
namespace=svc.namespace,
cost_category="total"
).set(svc.cost_per_million_requests)
cost_total_monthly.labels(
service=svc.service,
namespace=svc.namespace
).set(svc.total_cost)
Dashboarding: Making Costs Visible
The Service Cost Leaderboard
Every engineering team sees a weekly leaderboard showing their services ranked by CPM:
# Top 10 most expensive services by CPM
topk(10,
service_cost_per_million_requests{cost_category="total"}
)
# Services with CPM increasing week-over-week (cost regression)
service_cost_per_million_requests
- service_cost_per_million_requests offset 7d > 0
Cost Anomaly Alerts
We alert when a service's CPM increases significantly without a corresponding traffic change:
# alert-rules/cost-anomaly.yaml
groups:
- name: cost-anomalies
rules:
- alert: CostPerRequestSpike
expr: |
(
service_cost_per_million_requests
- service_cost_per_million_requests offset 7d
) / service_cost_per_million_requests offset 7d > 0.25
for: 1d
labels:
severity: warning
team: "{{ $labels.namespace }}"
annotations:
summary: "{{ $labels.service }} CPM increased >25% week-over-week"
dashboard: "https://grafana.internal/d/cost-per-request"
Optimization Discoveries
With CPM tracking live, we found immediate wins:
| Service | CPM Before | Issue Found | CPM After | Savings |
|---|---|---|---|---|
| auth-service | $4.82 | Redundant token validation calls | $1.23 | $8,400/mo |
| image-processor | $12.47 | Oversized pods (2x needed memory) | $5.11 | $6,200/mo |
| notification-service | $8.93 | N+1 database queries per batch | $2.14 | $11,800/mo |
| search-api | $6.31 | Cache miss rate 78% (config error) | $1.89 | $9,100/mo |
| analytics-ingest | $3.44 | Writing unqueried columns to DB | $1.92 | $5,500/mo |
Embedding Cost in Engineering Culture
PR Cost Impact Labels
Our CI pipeline estimates the cost impact of infrastructure changes:
#!/bin/bash
# scripts/estimate-cost-impact.sh
# Runs on PRs that modify Kubernetes manifests or Terraform
ADDED_CPU=$(diff_resource_requests cpu)
ADDED_MEMORY=$(diff_resource_requests memory)
# Estimate monthly cost impact
CPU_COST_PER_CORE=31.50 # Monthly cost per core in our cluster
MEM_COST_PER_GB=4.20 # Monthly cost per GB in our cluster
MONTHLY_IMPACT=$(echo "$ADDED_CPU * $CPU_COST_PER_CORE + $ADDED_MEMORY * $MEM_COST_PER_GB" | bc)
if (( $(echo "$MONTHLY_IMPACT > 100" | bc -l) )); then
gh pr comment --body "**Cost Impact**: +\$${MONTHLY_IMPACT}/month estimated. Please justify in PR description."
gh pr edit --add-label "cost-impact:high"
fi
Results After 6 Months
| Metric | Before | After |
|---|---|---|
| Average CPM across services | $5.84 | $2.21 |
| Monthly cloud spend | $183K | $142K |
| Cost attributed to services | 23% | 94% |
| Teams with cost visibility | 0 | 14 |
| Cost regressions caught in PR | 0 | 12/month |
| Optimization ROI (annualized) | — | $492K |
Key Takeaways
-
Cost per request is the only meaningful efficiency metric. Total cost conflates growth with waste. CPM isolates efficiency from scale.
-
Tag everything or accept blind spots. Untagged resources are unattributable costs. Enforce tagging at admission, not after the fact.
-
Shared cost allocation must be proportional. Equal splits punish efficient services and subsidize wasteful ones. Allocate by actual usage.
-
Make costs visible at the team level weekly. Engineers optimize what they measure. A weekly CPM leaderboard creates healthy competition.
-
Catch cost regressions in PR review. The cheapest time to fix a cost inefficiency is before it reaches production. Estimate impact during code review.
Cloud costs are engineering decisions with dollar signs. When every team knows their cost-per-request, optimization becomes everyone's job—not just FinOps.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.