Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Our ML platform runs 2,400 GPU-hours daily across training, fine-tuning, and inference workloads. When we measured the carbon footprint, it was sobering: 847 tonnes CO2e annually — equivalent to 184 cars driving for a year. After implementing carbon-aware scheduling, region-shifting for deferrable workloads, and inference optimization, we cut that to 491 tonnes without degrading model quality or serving latency. Here is the engineering behind carbon-aware AI infrastructure.
Why AI Carbon Emissions Matter Now
Large-scale AI is disproportionately carbon-intensive:
- A single GPT-4-class training run emits ~500 tonnes CO2e
- Inference at scale can exceed training emissions within months
- GPU instances consume 3-5x the power of equivalent CPU instances
- Most clouds run 24/7 regardless of grid carbon intensity
As AI workloads grow 4x year-over-year in most organizations, unchecked infrastructure emissions will become a regulatory and reputational risk. The EU Corporate Sustainability Reporting Directive (CSRD) now requires Scope 2 and 3 emissions disclosure for large companies — your cloud compute is Scope 2 (or Scope 3 if using a cloud provider).
Measuring: The Carbon Observability Stack
You cannot reduce what you cannot measure. We built a carbon observability pipeline that calculates real-time emissions per workload:
import httpx
import pandas as pd
from dataclasses import dataclass
from datetime import datetime
from typing import Optional
@dataclass
class CarbonIntensity:
region: str
timestamp: datetime
grams_co2_per_kwh: float
source: str # 'electricity_maps', 'watttime', 'cloud_provider'
renewable_percentage: float
@dataclass
class WorkloadEmission:
workload_id: str
workload_type: str # 'training', 'inference', 'fine-tuning'
region: str
gpu_type: str
gpu_hours: float
energy_kwh: float
carbon_grams: float
pue: float # Power Usage Effectiveness of the datacenter
start_time: datetime
end_time: datetime
class CarbonCalculator:
"""Calculate carbon emissions for GPU workloads using real-time grid data."""
# GPU TDP (Thermal Design Power) in watts
GPU_POWER = {
'a100-40gb': 250,
'a100-80gb': 300,
'h100-80gb': 700,
'l4': 72,
'l40s': 350,
't4': 70,
}
# Regional PUE estimates (source: cloud provider sustainability reports)
DATACENTER_PUE = {
'us-east-1': 1.15,
'us-west-2': 1.10,
'eu-west-1': 1.12,
'eu-north-1': 1.08, # Lowest PUE (cold climate)
'me-south-1': 1.35, # Highest PUE (hot climate)
}
def __init__(self, carbon_api_key: str):
self.api_key = carbon_api_key
self.intensity_cache: dict[str, CarbonIntensity] = {}
async def get_carbon_intensity(self, region: str) -> CarbonIntensity:
"""Fetch real-time carbon intensity for a cloud region."""
# Map AWS region to Electricity Maps zone
region_to_zone = {
'us-east-1': 'US-MIDA-PJM',
'us-west-2': 'US-NW-PACW',
'eu-west-1': 'IE',
'eu-north-1': 'SE-SE1',
'eu-central-1': 'DE',
'ap-northeast-1': 'JP-TK',
}
zone = region_to_zone.get(region, 'US-MIDA-PJM')
async with httpx.AsyncClient() as client:
response = await client.get(
f'https://api.electricitymap.org/v3/carbon-intensity/latest',
params={'zone': zone},
headers={'auth-token': self.api_key},
)
data = response.json()
return CarbonIntensity(
region=region,
timestamp=datetime.utcnow(),
grams_co2_per_kwh=data['carbonIntensity'],
source='electricity_maps',
renewable_percentage=data.get('renewablePercentage', 0),
)
def calculate_emission(self,
gpu_type: str,
gpu_count: int,
duration_hours: float,
region: str,
carbon_intensity: CarbonIntensity,
gpu_utilization: float = 0.8) -> WorkloadEmission:
"""Calculate carbon emissions for a GPU workload."""
# Step 1: Calculate energy consumption
tdp_watts = self.GPU_POWER.get(gpu_type, 300)
# Actual power = TDP * utilization factor (GPUs rarely run at 100% TDP)
actual_power_watts = tdp_watts * gpu_utilization
# Total energy including overhead (memory, networking, cooling)
overhead_factor = 1.15 # 15% overhead for supporting hardware
total_watts = actual_power_watts * gpu_count * overhead_factor
# Energy in kWh
energy_kwh = (total_watts / 1000) * duration_hours
# Step 2: Apply datacenter PUE (accounts for cooling, lighting, etc.)
pue = self.DATACENTER_PUE.get(region, 1.2)
effective_energy_kwh = energy_kwh * pue
# Step 3: Calculate carbon using grid intensity
carbon_grams = effective_energy_kwh * carbon_intensity.grams_co2_per_kwh
return WorkloadEmission(
workload_id='',
workload_type='',
region=region,
gpu_type=gpu_type,
gpu_hours=duration_hours * gpu_count,
energy_kwh=effective_energy_kwh,
carbon_grams=carbon_grams,
pue=pue,
start_time=datetime.utcnow(),
end_time=datetime.utcnow(),
)
Carbon-Aware Scheduling: Shifting Workloads in Time and Space
Not all AI workloads are latency-sensitive. Training jobs, batch inference, and fine-tuning can be deferred by hours or shifted to different regions without impacting users. We built a scheduler that exploits this flexibility:
from enum import Enum
class WorkloadPriority(Enum):
REAL_TIME = 'real_time' # Inference serving - cannot defer
NEAR_TIME = 'near_time' # Batch inference - 1-4 hour flexibility
DEFERRABLE = 'deferrable' # Training - 6-24 hour flexibility
@dataclass
class SchedulingDecision:
region: str
start_time: datetime
estimated_carbon_grams: float
carbon_savings_vs_immediate: float
delay_hours: float
class CarbonAwareScheduler:
"""Schedule GPU workloads based on carbon intensity forecasts."""
def __init__(self, calculator: CarbonCalculator, regions: list[str]):
self.calculator = calculator
self.available_regions = regions
async def find_optimal_schedule(self,
gpu_type: str,
gpu_count: int,
duration_hours: float,
priority: WorkloadPriority,
max_delay_hours: int = 24,
) -> SchedulingDecision:
"""Find the lowest-carbon time and region for a workload."""
# Define flexibility window based on priority
flexibility = {
WorkloadPriority.REAL_TIME: 0,
WorkloadPriority.NEAR_TIME: 4,
WorkloadPriority.DEFERRABLE: min(max_delay_hours, 24),
}
max_delay = flexibility[priority]
if max_delay == 0:
# Real-time: find lowest-carbon region right now
best_region = await self._find_greenest_region_now()
intensity = await self.calculator.get_carbon_intensity(best_region)
emission = self.calculator.calculate_emission(
gpu_type, gpu_count, duration_hours, best_region, intensity
)
return SchedulingDecision(
region=best_region,
start_time=datetime.utcnow(),
estimated_carbon_grams=emission.carbon_grams,
carbon_savings_vs_immediate=0,
delay_hours=0,
)
# For deferrable workloads: evaluate all region × time combinations
candidates = []
for region in self.available_regions:
# Get 24-hour carbon intensity forecast
forecast = await self._get_intensity_forecast(region, max_delay)
for hour_offset in range(0, max_delay):
avg_intensity = self._average_intensity(
forecast, hour_offset, duration_hours
)
emission = self.calculator.calculate_emission(
gpu_type, gpu_count, duration_hours, region,
CarbonIntensity(
region=region,
timestamp=datetime.utcnow(),
grams_co2_per_kwh=avg_intensity,
source='forecast',
renewable_percentage=0,
)
)
candidates.append(SchedulingDecision(
region=region,
start_time=datetime.utcnow() + pd.Timedelta(hours=hour_offset),
estimated_carbon_grams=emission.carbon_grams,
carbon_savings_vs_immediate=0, # Calculated below
delay_hours=hour_offset,
))
# Sort by carbon and pick the lowest
candidates.sort(key=lambda c: c.estimated_carbon_grams)
best = candidates[0]
immediate = next(
c for c in candidates if c.delay_hours == 0
)
best.carbon_savings_vs_immediate = (
immediate.estimated_carbon_grams - best.estimated_carbon_grams
)
return best
async def _find_greenest_region_now(self) -> str:
"""Find the region with lowest current carbon intensity."""
intensities = {}
for region in self.available_regions:
intensity = await self.calculator.get_carbon_intensity(region)
intensities[region] = intensity.grams_co2_per_kwh
return min(intensities, key=intensities.get)
async def _get_intensity_forecast(self, region: str, hours: int) -> list[float]:
"""Get hourly carbon intensity forecast for a region."""
# Production implementation calls Electricity Maps forecast API
# Returns list of gCO2/kWh values for each hour
return [200.0] * hours # Placeholder
def _average_intensity(self, forecast: list[float],
start_hour: int, duration: float) -> float:
"""Average carbon intensity over a workload's duration."""
end_hour = min(start_hour + int(duration) + 1, len(forecast))
relevant = forecast[start_hour:end_hour]
return sum(relevant) / len(relevant) if relevant else forecast[0]
Inference Optimization: Reducing Per-Request Carbon
For real-time inference that cannot be deferred, we reduce carbon by reducing compute:
| Optimization | Carbon Reduction | Latency Impact | Implementation |
|---|---|---|---|
| Model quantization (FP16 → INT8) | -45% GPU power | +3ms P99 | TensorRT / vLLM |
| Speculative decoding | -30% tokens generated | -15% TTFT | Custom implementation |
| Request batching (dynamic) | -35% per-request energy | +8ms avg | vLLM continuous batching |
| KV cache optimization | -20% memory power | None | PagedAttention |
| Spot instance mix for non-SLA traffic | -0% carbon, -60% cost | Variable | Karpenter provisioner |
# Karpenter provisioner for carbon-aware GPU scheduling
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: gpu-carbon-aware
spec:
template:
metadata:
labels:
workload-type: ai-inference
carbon-aware: "true"
spec:
requirements:
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["g", "p"] # GPU instance families
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: topology.kubernetes.io/zone
operator: In
values: ["us-west-2a", "us-west-2b"] # Hydro-powered zones
nodeClassRef:
name: gpu-al2-carbon
limits:
gpu: 64
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 5m
Results: 12-Month Carbon Reduction
| Source of Reduction | Annual Tonnes CO2e Saved | % of Total Reduction |
|---|---|---|
| Carbon-aware training scheduling | 148 | 42% |
| Region-shifting deferrable jobs | 89 | 25% |
| Inference model optimization | 67 | 19% |
| Hardware refresh (A100 → H100) | 34 | 10% |
| Idle GPU shutdown automation | 18 | 5% |
| Total | 356 | 100% |
From 847 tonnes/year to 491 tonnes/year — a 42% reduction.
The Business Case: Carbon as a Cost Proxy
Here is the insight that got executive buy-in: carbon reduction correlates with cost reduction at 0.87 R-squared. Lower energy means lower bills. Our carbon-aware scheduling saved $89K annually in electricity costs (reflected in cloud bills via usage-based GPU pricing).
| Metric | Before | After | Improvement |
|---|---|---|---|
| Annual CO2e emissions | 847 tonnes | 491 tonnes | -42% |
| Annual GPU compute cost | $1.2M | $980K | -18% |
| Average GPU utilization | 34% | 61% | +79% |
| Training jobs meeting SLA | 99.1% | 99.4% | +0.3% |
| Inference P99 latency | 142ms | 139ms | -2% (quantization gains) |
Carbon Reporting Dashboard
We publish a weekly carbon report to engineering leadership:
-- Weekly carbon report query
SELECT
DATE_TRUNC('week', start_time) as week,
workload_type,
team,
SUM(energy_kwh) as total_energy_kwh,
SUM(carbon_grams) / 1000 as total_carbon_kg,
AVG(carbon_grams / NULLIF(gpu_hours, 0)) as carbon_per_gpu_hour,
SUM(CASE WHEN was_deferred THEN carbon_savings_grams ELSE 0 END) / 1000
as deferred_savings_kg,
COUNT(CASE WHEN region != original_region THEN 1 END) as region_shifted_jobs
FROM workload_emissions
WHERE start_time > CURRENT_DATE - INTERVAL '7 days'
GROUP BY 1, 2, 3
ORDER BY total_carbon_kg DESC;
Lessons Learned
Carbon intensity varies 10x within a single day. In PJM (US East), carbon intensity ranges from 200 gCO2/kWh (midday solar) to 600 gCO2/kWh (evening peak). Shifting a 4-hour training job from 7 PM to 11 AM reduces emissions by 50-60% with zero performance impact.
Nordics are not always greenest. eu-north-1 (Stockholm) runs on 90%+ renewables, making it the obvious choice. But during Nordic winter evenings, carbon intensity spikes as fossil backup generation activates. Real-time data beats assumptions.
Hardware efficiency gains compound. H100 GPUs deliver 3x the compute per watt compared to A100. Migrating training workloads to newer hardware reduced both cost AND carbon, making it the easiest sell to leadership.
Do not sacrifice model quality for carbon. Aggressive quantization (INT4) reduced accuracy on our evaluation benchmarks by 2.3%. We found the sweet spot at INT8, which maintained accuracy within 0.1% while cutting inference energy by 45%.
Start with training, not inference. Training jobs are deferrable, high-energy, and infrequent — perfect candidates for carbon-aware scheduling. Inference optimization is harder because of latency constraints but delivers ongoing savings.
Conclusion
Carbon-aware AI infrastructure is not an altruistic choice — it is an engineering optimization that happens to align with sustainability goals. Reducing energy consumption reduces cost. Shifting workloads to renewable-heavy grids reduces both emissions and often price (renewable-heavy regions frequently have lower electricity costs). The tooling exists today: Electricity Maps for carbon data, Karpenter for region-aware scheduling, and quantization frameworks for inference optimization. The engineering challenge is building the observability pipeline and scheduling logic that makes carbon a first-class constraint alongside latency and cost. Start measuring, make deferrable workloads flexible, and let the data drive scheduling decisions.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning
How we used ML forecasting models to predict Black Friday traffic patterns, pre-provision infrastructure with surgical precision, and handle 47x normal load without wasting $180K on idle capacity.

Comments
No comments yet. Be the first to share your thoughts.