Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide

Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

#ai#sustainability#carbon-footprint#green-computing
Cover image for the article: Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide

Our ML platform runs 2,400 GPU-hours daily across training, fine-tuning, and inference workloads. When we measured the carbon footprint, it was sobering: 847 tonnes CO2e annually — equivalent to 184 cars driving for a year. After implementing carbon-aware scheduling, region-shifting for deferrable workloads, and inference optimization, we cut that to 491 tonnes without degrading model quality or serving latency. Here is the engineering behind carbon-aware AI infrastructure.

Why AI Carbon Emissions Matter Now

Large-scale AI is disproportionately carbon-intensive:

  • A single GPT-4-class training run emits ~500 tonnes CO2e
  • Inference at scale can exceed training emissions within months
  • GPU instances consume 3-5x the power of equivalent CPU instances
  • Most clouds run 24/7 regardless of grid carbon intensity

As AI workloads grow 4x year-over-year in most organizations, unchecked infrastructure emissions will become a regulatory and reputational risk. The EU Corporate Sustainability Reporting Directive (CSRD) now requires Scope 2 and 3 emissions disclosure for large companies — your cloud compute is Scope 2 (or Scope 3 if using a cloud provider).

Measuring: The Carbon Observability Stack

You cannot reduce what you cannot measure. We built a carbon observability pipeline that calculates real-time emissions per workload:

Carbon Observability Architecture

import httpx
import pandas as pd
from dataclasses import dataclass
from datetime import datetime
from typing import Optional

@dataclass
class CarbonIntensity:
    region: str
    timestamp: datetime
    grams_co2_per_kwh: float
    source: str  # 'electricity_maps', 'watttime', 'cloud_provider'
    renewable_percentage: float

@dataclass
class WorkloadEmission:
    workload_id: str
    workload_type: str  # 'training', 'inference', 'fine-tuning'
    region: str
    gpu_type: str
    gpu_hours: float
    energy_kwh: float
    carbon_grams: float
    pue: float  # Power Usage Effectiveness of the datacenter
    start_time: datetime
    end_time: datetime

class CarbonCalculator:
    """Calculate carbon emissions for GPU workloads using real-time grid data."""
    
    # GPU TDP (Thermal Design Power) in watts
    GPU_POWER = {
        'a100-40gb': 250,
        'a100-80gb': 300,
        'h100-80gb': 700,
        'l4': 72,
        'l40s': 350,
        't4': 70,
    }
    
    # Regional PUE estimates (source: cloud provider sustainability reports)
    DATACENTER_PUE = {
        'us-east-1': 1.15,
        'us-west-2': 1.10,
        'eu-west-1': 1.12,
        'eu-north-1': 1.08,  # Lowest PUE (cold climate)
        'me-south-1': 1.35,  # Highest PUE (hot climate)
    }
    
    def __init__(self, carbon_api_key: str):
        self.api_key = carbon_api_key
        self.intensity_cache: dict[str, CarbonIntensity] = {}
    
    async def get_carbon_intensity(self, region: str) -> CarbonIntensity:
        """Fetch real-time carbon intensity for a cloud region."""
        # Map AWS region to Electricity Maps zone
        region_to_zone = {
            'us-east-1': 'US-MIDA-PJM',
            'us-west-2': 'US-NW-PACW',
            'eu-west-1': 'IE',
            'eu-north-1': 'SE-SE1',
            'eu-central-1': 'DE',
            'ap-northeast-1': 'JP-TK',
        }
        
        zone = region_to_zone.get(region, 'US-MIDA-PJM')
        
        async with httpx.AsyncClient() as client:
            response = await client.get(
                f'https://api.electricitymap.org/v3/carbon-intensity/latest',
                params={'zone': zone},
                headers={'auth-token': self.api_key},
            )
            data = response.json()
        
        return CarbonIntensity(
            region=region,
            timestamp=datetime.utcnow(),
            grams_co2_per_kwh=data['carbonIntensity'],
            source='electricity_maps',
            renewable_percentage=data.get('renewablePercentage', 0),
        )
    
    def calculate_emission(self,
                           gpu_type: str,
                           gpu_count: int,
                           duration_hours: float,
                           region: str,
                           carbon_intensity: CarbonIntensity,
                           gpu_utilization: float = 0.8) -> WorkloadEmission:
        """Calculate carbon emissions for a GPU workload."""
        
        # Step 1: Calculate energy consumption
        tdp_watts = self.GPU_POWER.get(gpu_type, 300)
        # Actual power = TDP * utilization factor (GPUs rarely run at 100% TDP)
        actual_power_watts = tdp_watts * gpu_utilization
        
        # Total energy including overhead (memory, networking, cooling)
        overhead_factor = 1.15  # 15% overhead for supporting hardware
        total_watts = actual_power_watts * gpu_count * overhead_factor
        
        # Energy in kWh
        energy_kwh = (total_watts / 1000) * duration_hours
        
        # Step 2: Apply datacenter PUE (accounts for cooling, lighting, etc.)
        pue = self.DATACENTER_PUE.get(region, 1.2)
        effective_energy_kwh = energy_kwh * pue
        
        # Step 3: Calculate carbon using grid intensity
        carbon_grams = effective_energy_kwh * carbon_intensity.grams_co2_per_kwh
        
        return WorkloadEmission(
            workload_id='',
            workload_type='',
            region=region,
            gpu_type=gpu_type,
            gpu_hours=duration_hours * gpu_count,
            energy_kwh=effective_energy_kwh,
            carbon_grams=carbon_grams,
            pue=pue,
            start_time=datetime.utcnow(),
            end_time=datetime.utcnow(),
        )

Carbon-Aware Scheduling: Shifting Workloads in Time and Space

Not all AI workloads are latency-sensitive. Training jobs, batch inference, and fine-tuning can be deferred by hours or shifted to different regions without impacting users. We built a scheduler that exploits this flexibility:

from enum import Enum

class WorkloadPriority(Enum):
    REAL_TIME = 'real_time'     # Inference serving - cannot defer
    NEAR_TIME = 'near_time'    # Batch inference - 1-4 hour flexibility
    DEFERRABLE = 'deferrable'  # Training - 6-24 hour flexibility

@dataclass
class SchedulingDecision:
    region: str
    start_time: datetime
    estimated_carbon_grams: float
    carbon_savings_vs_immediate: float
    delay_hours: float

class CarbonAwareScheduler:
    """Schedule GPU workloads based on carbon intensity forecasts."""
    
    def __init__(self, calculator: CarbonCalculator, regions: list[str]):
        self.calculator = calculator
        self.available_regions = regions
    
    async def find_optimal_schedule(self,
                                    gpu_type: str,
                                    gpu_count: int,
                                    duration_hours: float,
                                    priority: WorkloadPriority,
                                    max_delay_hours: int = 24,
                                    ) -> SchedulingDecision:
        """Find the lowest-carbon time and region for a workload."""
        
        # Define flexibility window based on priority
        flexibility = {
            WorkloadPriority.REAL_TIME: 0,
            WorkloadPriority.NEAR_TIME: 4,
            WorkloadPriority.DEFERRABLE: min(max_delay_hours, 24),
        }
        
        max_delay = flexibility[priority]
        
        if max_delay == 0:
            # Real-time: find lowest-carbon region right now
            best_region = await self._find_greenest_region_now()
            intensity = await self.calculator.get_carbon_intensity(best_region)
            emission = self.calculator.calculate_emission(
                gpu_type, gpu_count, duration_hours, best_region, intensity
            )
            return SchedulingDecision(
                region=best_region,
                start_time=datetime.utcnow(),
                estimated_carbon_grams=emission.carbon_grams,
                carbon_savings_vs_immediate=0,
                delay_hours=0,
            )
        
        # For deferrable workloads: evaluate all region × time combinations
        candidates = []
        
        for region in self.available_regions:
            # Get 24-hour carbon intensity forecast
            forecast = await self._get_intensity_forecast(region, max_delay)
            
            for hour_offset in range(0, max_delay):
                avg_intensity = self._average_intensity(
                    forecast, hour_offset, duration_hours
                )
                
                emission = self.calculator.calculate_emission(
                    gpu_type, gpu_count, duration_hours, region,
                    CarbonIntensity(
                        region=region,
                        timestamp=datetime.utcnow(),
                        grams_co2_per_kwh=avg_intensity,
                        source='forecast',
                        renewable_percentage=0,
                    )
                )
                
                candidates.append(SchedulingDecision(
                    region=region,
                    start_time=datetime.utcnow() + pd.Timedelta(hours=hour_offset),
                    estimated_carbon_grams=emission.carbon_grams,
                    carbon_savings_vs_immediate=0,  # Calculated below
                    delay_hours=hour_offset,
                ))
        
        # Sort by carbon and pick the lowest
        candidates.sort(key=lambda c: c.estimated_carbon_grams)
        best = candidates[0]
        immediate = next(
            c for c in candidates if c.delay_hours == 0
        )
        best.carbon_savings_vs_immediate = (
            immediate.estimated_carbon_grams - best.estimated_carbon_grams
        )
        
        return best
    
    async def _find_greenest_region_now(self) -> str:
        """Find the region with lowest current carbon intensity."""
        intensities = {}
        for region in self.available_regions:
            intensity = await self.calculator.get_carbon_intensity(region)
            intensities[region] = intensity.grams_co2_per_kwh
        return min(intensities, key=intensities.get)
    
    async def _get_intensity_forecast(self, region: str, hours: int) -> list[float]:
        """Get hourly carbon intensity forecast for a region."""
        # Production implementation calls Electricity Maps forecast API
        # Returns list of gCO2/kWh values for each hour
        return [200.0] * hours  # Placeholder
    
    def _average_intensity(self, forecast: list[float], 
                           start_hour: int, duration: float) -> float:
        """Average carbon intensity over a workload's duration."""
        end_hour = min(start_hour + int(duration) + 1, len(forecast))
        relevant = forecast[start_hour:end_hour]
        return sum(relevant) / len(relevant) if relevant else forecast[0]

Inference Optimization: Reducing Per-Request Carbon

For real-time inference that cannot be deferred, we reduce carbon by reducing compute:

OptimizationCarbon ReductionLatency ImpactImplementation
Model quantization (FP16 → INT8)-45% GPU power+3ms P99TensorRT / vLLM
Speculative decoding-30% tokens generated-15% TTFTCustom implementation
Request batching (dynamic)-35% per-request energy+8ms avgvLLM continuous batching
KV cache optimization-20% memory powerNonePagedAttention
Spot instance mix for non-SLA traffic-0% carbon, -60% costVariableKarpenter provisioner
# Karpenter provisioner for carbon-aware GPU scheduling
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
  name: gpu-carbon-aware
spec:
  template:
    metadata:
      labels:
        workload-type: ai-inference
        carbon-aware: "true"
    spec:
      requirements:
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["g", "p"]  # GPU instance families
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]
        - key: topology.kubernetes.io/zone
          operator: In
          values: ["us-west-2a", "us-west-2b"]  # Hydro-powered zones
      nodeClassRef:
        name: gpu-al2-carbon
  limits:
    gpu: 64
  disruption:
    consolidationPolicy: WhenUnderutilized
    consolidateAfter: 5m

Results: 12-Month Carbon Reduction

Source of ReductionAnnual Tonnes CO2e Saved% of Total Reduction
Carbon-aware training scheduling14842%
Region-shifting deferrable jobs8925%
Inference model optimization6719%
Hardware refresh (A100 → H100)3410%
Idle GPU shutdown automation185%
Total356100%

From 847 tonnes/year to 491 tonnes/year — a 42% reduction.

Carbon Reduction Timeline

The Business Case: Carbon as a Cost Proxy

Here is the insight that got executive buy-in: carbon reduction correlates with cost reduction at 0.87 R-squared. Lower energy means lower bills. Our carbon-aware scheduling saved $89K annually in electricity costs (reflected in cloud bills via usage-based GPU pricing).

MetricBeforeAfterImprovement
Annual CO2e emissions847 tonnes491 tonnes-42%
Annual GPU compute cost$1.2M$980K-18%
Average GPU utilization34%61%+79%
Training jobs meeting SLA99.1%99.4%+0.3%
Inference P99 latency142ms139ms-2% (quantization gains)

Carbon Reporting Dashboard

We publish a weekly carbon report to engineering leadership:

-- Weekly carbon report query
SELECT
  DATE_TRUNC('week', start_time) as week,
  workload_type,
  team,
  SUM(energy_kwh) as total_energy_kwh,
  SUM(carbon_grams) / 1000 as total_carbon_kg,
  AVG(carbon_grams / NULLIF(gpu_hours, 0)) as carbon_per_gpu_hour,
  SUM(CASE WHEN was_deferred THEN carbon_savings_grams ELSE 0 END) / 1000 
    as deferred_savings_kg,
  COUNT(CASE WHEN region != original_region THEN 1 END) as region_shifted_jobs
FROM workload_emissions
WHERE start_time > CURRENT_DATE - INTERVAL '7 days'
GROUP BY 1, 2, 3
ORDER BY total_carbon_kg DESC;

Lessons Learned

Carbon intensity varies 10x within a single day. In PJM (US East), carbon intensity ranges from 200 gCO2/kWh (midday solar) to 600 gCO2/kWh (evening peak). Shifting a 4-hour training job from 7 PM to 11 AM reduces emissions by 50-60% with zero performance impact.

Nordics are not always greenest. eu-north-1 (Stockholm) runs on 90%+ renewables, making it the obvious choice. But during Nordic winter evenings, carbon intensity spikes as fossil backup generation activates. Real-time data beats assumptions.

Hardware efficiency gains compound. H100 GPUs deliver 3x the compute per watt compared to A100. Migrating training workloads to newer hardware reduced both cost AND carbon, making it the easiest sell to leadership.

Do not sacrifice model quality for carbon. Aggressive quantization (INT4) reduced accuracy on our evaluation benchmarks by 2.3%. We found the sweet spot at INT8, which maintained accuracy within 0.1% while cutting inference energy by 45%.

Start with training, not inference. Training jobs are deferrable, high-energy, and infrequent — perfect candidates for carbon-aware scheduling. Inference optimization is harder because of latency constraints but delivers ongoing savings.

Conclusion

Carbon-aware AI infrastructure is not an altruistic choice — it is an engineering optimization that happens to align with sustainability goals. Reducing energy consumption reduces cost. Shifting workloads to renewable-heavy grids reduces both emissions and often price (renewable-heavy regions frequently have lower electricity costs). The tooling exists today: Electricity Maps for carbon data, Karpenter for region-aware scheduling, and quantization frameworks for inference optimization. The engineering challenge is building the observability pipeline and scheduling logic that makes carbon a first-class constraint alongside latency and cost. Start measuring, make deferrable workloads flexible, and let the data drive scheduling decisions.

Comments

    No comments yet. Be the first to share your thoughts.