Tracking Cost-Per-Request Across Every Microservice

How to implement unit economics tracking for cloud infrastructure, allocating real costs to individual services and API endpoints to drive optimization decisions.

#cost-optimization#unit-economics#finops#metrics
Cover image for the article: Tracking Cost-Per-Request Across Every Microservice

The Problem: Cloud Costs Without Attribution

Our monthly AWS bill grew from $47K to $183K in 18 months. Everyone agreed costs were too high, but nobody could answer the basic question: which services cost the most per request, and why?

Without unit economics, cost optimization is guesswork. Teams optimize what's visible (that one big EC2 instance) while ignoring systemic inefficiencies (a service making 47 redundant database calls per request). We needed a system that attributes every dollar of cloud spend to the requests it serves.

After implementing cost-per-request tracking, we identified $41K/month in optimization opportunities within the first 30 days. Six months later, our cost-per-million-requests dropped 62% while traffic grew 40%.

The Unit Economics Framework

The core metric: cost per 1 million requests (CPM) for each service. This normalizes cost against traffic and makes services comparable regardless of scale.

Cost Per Request Architecture

CPM = (Monthly Infrastructure Cost for Service) / (Monthly Requests / 1,000,000)

But "infrastructure cost for service" isn't straightforward in shared environments. We break it into three allocation layers:

  1. Direct costs: Resources exclusively used by one service (dedicated pods, databases)
  2. Shared costs: Resources used by multiple services (load balancers, shared caches, ingress)
  3. Platform costs: Cluster overhead (control plane, monitoring, networking)

Implementation: Cost Allocation Pipeline

Step 1: Tag Everything

Every resource must be attributed to a service. We enforce tagging through admission webhooks:

# admission-webhook/tag-enforcement.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingWebhookConfiguration
metadata:
  name: cost-allocation-tagger
webhooks:
  - name: cost-tagger.finops.internal
    rules:
      - apiGroups: ["apps"]
        resources: ["deployments", "statefulsets"]
        operations: ["CREATE", "UPDATE"]
    clientConfig:
      service:
        name: cost-tagger
        namespace: finops
---
# The webhook injects these labels if missing:
# app.kubernetes.io/name: <service-name>
# finops.internal/cost-center: <team>
# finops.internal/environment: <env>

Step 2: Collect Resource Usage Per Service

We use Kubecost's allocation API combined with Prometheus metrics:

# finops/cost_collector.py
from dataclasses import dataclass
from datetime import datetime

@dataclass
class ServiceCost:
    service: str
    namespace: str
    period: str  # e.g., "2026-08"
    
    # Direct compute costs
    cpu_cost: float  # Based on actual usage, not requests
    memory_cost: float
    gpu_cost: float
    
    # Storage costs
    pv_cost: float
    s3_cost: float
    
    # Network costs
    egress_cost: float
    lb_cost: float
    
    # Shared cost allocation
    shared_infra_cost: float
    platform_overhead_cost: float
    
    # Traffic
    total_requests: int
    
    @property
    def total_cost(self) -> float:
        return (
            self.cpu_cost + self.memory_cost + self.gpu_cost +
            self.pv_cost + self.s3_cost +
            self.egress_cost + self.lb_cost +
            self.shared_infra_cost + self.platform_overhead_cost
        )
    
    @property
    def cost_per_million_requests(self) -> float:
        if self.total_requests == 0:
            return 0
        return (self.total_cost / self.total_requests) * 1_000_000
    
    @property
    def cost_breakdown(self) -> dict[str, float]:
        """Percentage breakdown by cost category."""
        total = self.total_cost
        if total == 0:
            return {}
        return {
            "compute": (self.cpu_cost + self.memory_cost + self.gpu_cost) / total * 100,
            "storage": (self.pv_cost + self.s3_cost) / total * 100,
            "network": (self.egress_cost + self.lb_cost) / total * 100,
            "shared": (self.shared_infra_cost + self.platform_overhead_cost) / total * 100,
        }

Step 3: Shared Cost Allocation

Shared resources (ALBs, NAT gateways, monitoring infrastructure) are allocated proportionally based on usage:

# finops/shared_allocation.py
def allocate_shared_costs(
    shared_resource_cost: float,
    resource_type: str,
    service_usage: dict[str, float],
) -> dict[str, float]:
    """
    Allocate shared costs proportionally based on usage metrics.
    
    Allocation strategies by resource type:
    - load_balancer: by request count
    - nat_gateway: by egress bytes  
    - monitoring: by time series count
    - platform: by cpu_seconds consumed
    """
    total_usage = sum(service_usage.values())
    if total_usage == 0:
        # Equal split if no usage data
        equal_share = shared_resource_cost / len(service_usage)
        return {svc: equal_share for svc in service_usage}
    
    return {
        service: shared_resource_cost * (usage / total_usage)
        for service, usage in service_usage.items()
    }

Step 4: Prometheus Metrics Export

We export CPM as a Prometheus metric for dashboarding and alerting:

# finops/metrics_exporter.py
from prometheus_client import Gauge, start_http_server

cost_per_million = Gauge(
    'service_cost_per_million_requests',
    'Cost in USD per 1M requests',
    ['service', 'namespace', 'cost_category']
)

cost_total_monthly = Gauge(
    'service_cost_total_monthly_usd',
    'Total monthly cost in USD',
    ['service', 'namespace']
)

cost_efficiency_score = Gauge(
    'service_cost_efficiency_score',
    'Efficiency score (0-100, higher is better)',
    ['service', 'namespace']
)

def update_metrics(services: list[ServiceCost]):
    for svc in services:
        cost_per_million.labels(
            service=svc.service,
            namespace=svc.namespace,
            cost_category="total"
        ).set(svc.cost_per_million_requests)
        
        cost_total_monthly.labels(
            service=svc.service,
            namespace=svc.namespace
        ).set(svc.total_cost)

Dashboarding: Making Costs Visible

The Service Cost Leaderboard

Every engineering team sees a weekly leaderboard showing their services ranked by CPM:

# Top 10 most expensive services by CPM
topk(10,
  service_cost_per_million_requests{cost_category="total"}
)

# Services with CPM increasing week-over-week (cost regression)
service_cost_per_million_requests 
  - service_cost_per_million_requests offset 7d > 0

Cost Per Request Dashboard

Cost Anomaly Alerts

We alert when a service's CPM increases significantly without a corresponding traffic change:

# alert-rules/cost-anomaly.yaml
groups:
  - name: cost-anomalies
    rules:
      - alert: CostPerRequestSpike
        expr: |
          (
            service_cost_per_million_requests 
            - service_cost_per_million_requests offset 7d
          ) / service_cost_per_million_requests offset 7d > 0.25
        for: 1d
        labels:
          severity: warning
          team: "{{ $labels.namespace }}"
        annotations:
          summary: "{{ $labels.service }} CPM increased >25% week-over-week"
          dashboard: "https://grafana.internal/d/cost-per-request"

Optimization Discoveries

With CPM tracking live, we found immediate wins:

ServiceCPM BeforeIssue FoundCPM AfterSavings
auth-service$4.82Redundant token validation calls$1.23$8,400/mo
image-processor$12.47Oversized pods (2x needed memory)$5.11$6,200/mo
notification-service$8.93N+1 database queries per batch$2.14$11,800/mo
search-api$6.31Cache miss rate 78% (config error)$1.89$9,100/mo
analytics-ingest$3.44Writing unqueried columns to DB$1.92$5,500/mo

Embedding Cost in Engineering Culture

PR Cost Impact Labels

Our CI pipeline estimates the cost impact of infrastructure changes:

#!/bin/bash
# scripts/estimate-cost-impact.sh
# Runs on PRs that modify Kubernetes manifests or Terraform

ADDED_CPU=$(diff_resource_requests cpu)
ADDED_MEMORY=$(diff_resource_requests memory)

# Estimate monthly cost impact
CPU_COST_PER_CORE=31.50  # Monthly cost per core in our cluster
MEM_COST_PER_GB=4.20     # Monthly cost per GB in our cluster

MONTHLY_IMPACT=$(echo "$ADDED_CPU * $CPU_COST_PER_CORE + $ADDED_MEMORY * $MEM_COST_PER_GB" | bc)

if (( $(echo "$MONTHLY_IMPACT > 100" | bc -l) )); then
  gh pr comment --body "**Cost Impact**: +\$${MONTHLY_IMPACT}/month estimated. Please justify in PR description."
  gh pr edit --add-label "cost-impact:high"
fi

Results After 6 Months

MetricBeforeAfter
Average CPM across services$5.84$2.21
Monthly cloud spend$183K$142K
Cost attributed to services23%94%
Teams with cost visibility014
Cost regressions caught in PR012/month
Optimization ROI (annualized)—$492K

Key Takeaways

  1. Cost per request is the only meaningful efficiency metric. Total cost conflates growth with waste. CPM isolates efficiency from scale.

  2. Tag everything or accept blind spots. Untagged resources are unattributable costs. Enforce tagging at admission, not after the fact.

  3. Shared cost allocation must be proportional. Equal splits punish efficient services and subsidize wasteful ones. Allocate by actual usage.

  4. Make costs visible at the team level weekly. Engineers optimize what they measure. A weekly CPM leaderboard creates healthy competition.

  5. Catch cost regressions in PR review. The cheapest time to fix a cost inefficiency is before it reaches production. Estimate impact during code review.

Cloud costs are engineering decisions with dollar signs. When every team knows their cost-per-request, optimization becomes everyone's job—not just FinOps.

Comments

    No comments yet. Be the first to share your thoughts.