AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning

How we used ML forecasting models to predict Black Friday traffic patterns, pre-provision infrastructure with surgical precision, and handle 47x normal load without wasting $180K on idle capacity.

#capacity-planning#ai#forecasting#infrastructure
Cover image for the article: AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning

Black Friday 2025 hit our platform with 47x normal traffic — 680,000 concurrent users compared to our typical 14,500. The year before, we handled it by over-provisioning everything 60x, burning $180K in pre-staged infrastructure that sat 70% idle even at peak. This year, our AI-driven capacity planning system predicted the load curve within 6% accuracy, pre-provisioned exactly what we needed, and we spent $52K total. Same uptime SLA (99.99%), $128K less. Here is the complete system.

The Traditional Approach: Spreadsheet-Driven Terror

Every year, the same ritual: engineering leadership sits in a room with last year's metrics, applies a growth multiplier, adds a "safety factor" (usually 2x because nobody wants to be the person who under-provisioned), and submits a purchase order. The result:

  • 2023 Black Friday: Provisioned 80x normal. Actual peak: 31x. Utilization: 39%.
  • 2024 Black Friday: Provisioned 60x normal. Actual peak: 38x. Utilization: 63%.
  • 2025 target: Provision exactly what we need, with intelligent buffering.

The problem with manual capacity planning is that it optimizes for one metric: "did we stay up?" But the question should be: "did we stay up efficiently?"

Architecture: Multi-Signal Forecasting Pipeline

Capacity Planning Architecture

Our capacity planning system ingests four signal categories:

  1. Historical traffic patterns — 3 years of Black Friday data at 1-minute granularity
  2. Business signals — Marketing campaign schedules, promotional discount depths, email send volumes
  3. External indicators — Social media mention velocity, competitor promotional timing, macroeconomic sentiment
  4. Real-time telemetry — Current-day traffic shape for last-mile adjustments
import numpy as np
import pandas as pd
from dataclasses import dataclass
from typing import Optional
from sklearn.ensemble import RandomForestRegressor
from scipy.stats import norm

@dataclass
class CapacityForecast:
    timestamp: pd.Timestamp
    predicted_rps: float
    p50_rps: float
    p75_rps: float
    p95_rps: float
    p99_rps: float
    required_instances: dict  # service_name -> instance_count
    estimated_cost: float
    confidence: float

class BlackFridayCapacityModel:
    """Multi-signal model for high-traffic event capacity planning."""
    
    def __init__(self, services_config: dict):
        self.services = services_config  # Per-service capacity specs
        self.models = {}  # Per-service trained models
        self.historical_peaks = []
    
    def train(self, 
              historical_events: list[pd.DataFrame],
              business_context: pd.DataFrame) -> dict:
        """
        Train on historical high-traffic events.
        
        Args:
            historical_events: List of DataFrames, one per past event,
                             with columns [timestamp, rps, service, error_rate]
            business_context: DataFrame with promotional/marketing signals
        """
        metrics = {}
        
        for service_name in self.services:
            # Combine all historical events for this service
            service_data = pd.concat([
                df[df['service'] == service_name] 
                for df in historical_events
            ])
            
            # Feature engineering
            features = self._build_event_features(service_data, business_context)
            targets = service_data['rps'].values
            
            # Train ensemble model
            model = RandomForestRegressor(
                n_estimators=500,
                max_depth=12,
                min_samples_leaf=10,
                n_jobs=-1,
            )
            model.fit(features, targets)
            self.models[service_name] = model
            
            # Calculate prediction intervals using OOB residuals
            oob_predictions = model.oob_prediction_
            residuals = targets - oob_predictions
            metrics[service_name] = {
                'mape': np.mean(np.abs(residuals / targets)) * 100,
                'residual_std': np.std(residuals),
            }
        
        return metrics
    
    def forecast_event(self,
                       event_date: pd.Timestamp,
                       business_signals: dict,
                       duration_hours: int = 24) -> list[CapacityForecast]:
        """Generate minute-by-minute capacity forecast for the event."""
        
        forecasts = []
        timestamps = pd.date_range(
            start=event_date,
            periods=duration_hours * 60,
            freq='1min',
        )
        
        for ts in timestamps:
            service_instances = {}
            total_cost = 0.0
            total_confidence = 0.0
            
            for service_name, config in self.services.items():
                model = self.models[service_name]
                features = self._build_prediction_features(
                    ts, business_signals
                )
                
                # Point prediction
                predicted_rps = model.predict(features.reshape(1, -1))[0]
                
                # Prediction intervals via quantile estimation
                tree_predictions = np.array([
                    tree.predict(features.reshape(1, -1))[0]
                    for tree in model.estimators_
                ])
                
                p50 = np.percentile(tree_predictions, 50)
                p75 = np.percentile(tree_predictions, 75)
                p95 = np.percentile(tree_predictions, 95)
                p99 = np.percentile(tree_predictions, 99)
                
                # Convert RPS to instance count
                capacity_per_instance = config['rps_capacity']
                # Use P95 for provisioning (handles most scenarios)
                instances_needed = int(np.ceil(
                    p95 / capacity_per_instance * 1.2  # 20% buffer
                ))
                instances_needed = max(instances_needed, config['min_instances'])
                
                service_instances[service_name] = instances_needed
                total_cost += instances_needed * config['hourly_cost'] / 60
                
                # Confidence = inverse of coefficient of variation
                cv = np.std(tree_predictions) / np.mean(tree_predictions)
                total_confidence += max(0, 1 - cv)
            
            avg_confidence = total_confidence / len(self.services)
            
            forecasts.append(CapacityForecast(
                timestamp=ts,
                predicted_rps=predicted_rps,
                p50_rps=p50,
                p75_rps=p75,
                p95_rps=p95,
                p99_rps=p99,
                required_instances=service_instances,
                estimated_cost=total_cost,
                confidence=avg_confidence,
            ))
        
        return forecasts
    
    def _build_event_features(self, data, context) -> np.ndarray:
        """Build feature matrix from historical event data."""
        ts = pd.DatetimeIndex(data['timestamp'])
        features = pd.DataFrame({
            'hour': ts.hour,
            'minute': ts.minute,
            'minutes_since_midnight': ts.hour * 60 + ts.minute,
            'is_peak_window': ((ts.hour >= 8) & (ts.hour <= 22)).astype(int),
            'hour_sin': np.sin(2 * np.pi * ts.hour / 24),
            'hour_cos': np.cos(2 * np.pi * ts.hour / 24),
            'year_growth': (ts.year - 2022).values,  # YoY growth factor
        })
        return features.values
    
    def _build_prediction_features(self, ts, signals) -> np.ndarray:
        """Build features for a single prediction timestamp."""
        return np.array([
            ts.hour,
            ts.minute,
            ts.hour * 60 + ts.minute,
            1 if 8 <= ts.hour <= 22 else 0,
            np.sin(2 * np.pi * ts.hour / 24),
            np.cos(2 * np.pi * ts.hour / 24),
            ts.year - 2022,
        ])

The Provisioning Plan Generator

The forecast becomes an actionable provisioning timeline — exactly when to scale each service and to what level:

@dataclass
class ProvisioningAction:
    timestamp: pd.Timestamp
    service: str
    action: str  # 'scale_up', 'scale_down', 'pre_warm'
    target_instances: int
    current_instances: int
    lead_time_minutes: int
    estimated_cost_delta: float

def generate_provisioning_plan(
    forecasts: list[CapacityForecast],
    current_state: dict,
    boot_time_minutes: int = 5,
) -> list[ProvisioningAction]:
    """Convert forecasts into a time-ordered provisioning plan."""
    
    actions = []
    previous_targets = {svc: current_state[svc] for svc in current_state}
    
    for forecast in forecasts:
        for service, target in forecast.required_instances.items():
            current = previous_targets.get(service, 0)
            
            # Only generate action if change exceeds 10% threshold
            if abs(target - current) / max(current, 1) < 0.1:
                continue
            
            if target > current:
                # Scale up: schedule BEFORE the need (lead time)
                action_time = forecast.timestamp - pd.Timedelta(
                    minutes=boot_time_minutes + 2  # Boot + health check
                )
                actions.append(ProvisioningAction(
                    timestamp=action_time,
                    service=service,
                    action='scale_up',
                    target_instances=target,
                    current_instances=current,
                    lead_time_minutes=boot_time_minutes,
                    estimated_cost_delta=(target - current) * 0.15,  # $/hr per instance
                ))
            elif target < current and current - target > 5:
                # Scale down: only if significant reduction, with delay
                action_time = forecast.timestamp + pd.Timedelta(minutes=15)
                actions.append(ProvisioningAction(
                    timestamp=action_time,
                    service=service,
                    action='scale_down',
                    target_instances=target,
                    current_instances=current,
                    lead_time_minutes=0,
                    estimated_cost_delta=(target - current) * 0.15,
                ))
            
            previous_targets[service] = target
    
    # Deduplicate: keep only the last action per service per 10-min window
    return _deduplicate_actions(actions)

Black Friday 2025: Actual vs Predicted

The model generated a provisioning timeline with 1,440 data points (one per minute for 24 hours). Here is how it performed:

ServicePredicted Peak RPSActual Peak RPSProvisioned InstancesPeak Utilization
API Gateway142,000134,8008478%
Order Service28,00031,2004286%
Payment Service18,00016,4002472%
Search Service89,00095,2006291%
Product Catalog210,000198,0004882%

The Search Service was the only near-miss: 91% utilization at peak, dangerously close to our 85% comfort threshold. The model underestimated search traffic because this year's promotional strategy featured a new "search and win" gamification that was not present in historical data.

Predicted vs Actual Traffic

Cost Comparison: AI-Driven vs Manual

ApproachTotal Event CostPeak UtilizationOver-Provisioned By
2023 (manual, 80x)$210K39%105%
2024 (manual, 60x)$180K63%58%
2025 (AI-driven)$52K82% avg18%

The $128K savings came from two sources:

  1. Right-sized peak: Provisioned 47x instead of 60x ($82K saved)
  2. Shaped provisioning curve: Scaled up gradually and down after peak instead of flat-topping for 24 hours ($46K saved)

The Feedback Loop: Real-Time Adjustments

Even with a 30-day-ahead forecast, we run real-time adjustments on the day of the event. A lightweight model compares the morning's actual traffic shape to the predicted shape and adjusts afternoon forecasts:

def real_time_adjustment(
    forecast: list[CapacityForecast],
    actual_so_far: pd.DataFrame,
    adjustment_timestamp: pd.Timestamp,
) -> list[CapacityForecast]:
    """Adjust remaining forecast based on observed vs predicted ratio."""
    
    # Compare last 2 hours of actual vs predicted
    lookback = adjustment_timestamp - pd.Timedelta(hours=2)
    recent_actual = actual_so_far[actual_so_far['timestamp'] > lookback]
    recent_forecast = [
        f for f in forecast 
        if lookback < f.timestamp <= adjustment_timestamp
    ]
    
    if not recent_forecast:
        return forecast
    
    # Calculate scaling factor
    actual_mean = recent_actual['rps'].mean()
    predicted_mean = np.mean([f.predicted_rps for f in recent_forecast])
    
    if predicted_mean == 0:
        return forecast
    
    scale_factor = actual_mean / predicted_mean
    
    # Apply dampened adjustment to future predictions
    # Dampen to avoid overreacting to short-term variance
    dampened_factor = 1.0 + (scale_factor - 1.0) * 0.6
    
    adjusted = []
    for f in forecast:
        if f.timestamp > adjustment_timestamp:
            adjusted.append(CapacityForecast(
                timestamp=f.timestamp,
                predicted_rps=f.predicted_rps * dampened_factor,
                p50_rps=f.p50_rps * dampened_factor,
                p75_rps=f.p75_rps * dampened_factor,
                p95_rps=f.p95_rps * dampened_factor,
                p99_rps=f.p99_rps * dampened_factor,
                required_instances={
                    svc: int(np.ceil(count * dampened_factor))
                    for svc, count in f.required_instances.items()
                },
                estimated_cost=f.estimated_cost * dampened_factor,
                confidence=f.confidence * 0.9,  # Slight confidence reduction
            ))
        else:
            adjusted.append(f)
    
    return adjusted

On Black Friday morning, the model detected that actual traffic was running 8% above forecast. The real-time adjustment bumped afternoon predictions up by ~5% (dampened), which triggered an additional 6 instances across services 45 minutes before the actual peak. Those 6 instances prevented what would have been a search service saturation event.

Lessons Learned

Three years of event data is the minimum. One year gives you a shape, two years give you a trend, three years give you confidence intervals. We could not build this system before 2025 because we only started collecting minute-level event data in 2022.

Business signals matter more than historical patterns. A 20% deeper promotional discount drives 35% more traffic. A new product category launch changes the traffic mix entirely. Feed your model the marketing plan, not just last year's metrics.

Test the provisioning plan with chaos engineering. Two weeks before Black Friday, we ran a simulated load test at 50x using the predicted traffic shape. This validated both the model's output and our infrastructure's ability to handle the predicted load.

Scale down conservatively. Our model aggressively scaled down after midnight, which was correct in 2024 but wrong in 2025 — international customers in different timezones created a longer tail. We added a 2-hour delay buffer for scale-down actions.

Conclusion

AI-assisted capacity planning transforms peak traffic events from existential threats into engineering problems with quantifiable solutions. The model is not magic — it is time-series forecasting with business context injection and real-time feedback. The value comes from replacing human intuition ("let's do 60x to be safe") with data-driven precision ("47x with shaped provisioning handles P95 scenarios"). For any organization spending more than $50K on peak-event infrastructure, the ML investment pays for itself in a single event cycle.

Comments

    No comments yet. Be the first to share your thoughts.