AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning
How we used ML forecasting models to predict Black Friday traffic patterns, pre-provision infrastructure with surgical precision, and handle 47x normal load without wasting $180K on idle capacity.

Black Friday 2025 hit our platform with 47x normal traffic — 680,000 concurrent users compared to our typical 14,500. The year before, we handled it by over-provisioning everything 60x, burning $180K in pre-staged infrastructure that sat 70% idle even at peak. This year, our AI-driven capacity planning system predicted the load curve within 6% accuracy, pre-provisioned exactly what we needed, and we spent $52K total. Same uptime SLA (99.99%), $128K less. Here is the complete system.
The Traditional Approach: Spreadsheet-Driven Terror
Every year, the same ritual: engineering leadership sits in a room with last year's metrics, applies a growth multiplier, adds a "safety factor" (usually 2x because nobody wants to be the person who under-provisioned), and submits a purchase order. The result:
- 2023 Black Friday: Provisioned 80x normal. Actual peak: 31x. Utilization: 39%.
- 2024 Black Friday: Provisioned 60x normal. Actual peak: 38x. Utilization: 63%.
- 2025 target: Provision exactly what we need, with intelligent buffering.
The problem with manual capacity planning is that it optimizes for one metric: "did we stay up?" But the question should be: "did we stay up efficiently?"
Architecture: Multi-Signal Forecasting Pipeline
Our capacity planning system ingests four signal categories:
- Historical traffic patterns — 3 years of Black Friday data at 1-minute granularity
- Business signals — Marketing campaign schedules, promotional discount depths, email send volumes
- External indicators — Social media mention velocity, competitor promotional timing, macroeconomic sentiment
- Real-time telemetry — Current-day traffic shape for last-mile adjustments
import numpy as np
import pandas as pd
from dataclasses import dataclass
from typing import Optional
from sklearn.ensemble import RandomForestRegressor
from scipy.stats import norm
@dataclass
class CapacityForecast:
timestamp: pd.Timestamp
predicted_rps: float
p50_rps: float
p75_rps: float
p95_rps: float
p99_rps: float
required_instances: dict # service_name -> instance_count
estimated_cost: float
confidence: float
class BlackFridayCapacityModel:
"""Multi-signal model for high-traffic event capacity planning."""
def __init__(self, services_config: dict):
self.services = services_config # Per-service capacity specs
self.models = {} # Per-service trained models
self.historical_peaks = []
def train(self,
historical_events: list[pd.DataFrame],
business_context: pd.DataFrame) -> dict:
"""
Train on historical high-traffic events.
Args:
historical_events: List of DataFrames, one per past event,
with columns [timestamp, rps, service, error_rate]
business_context: DataFrame with promotional/marketing signals
"""
metrics = {}
for service_name in self.services:
# Combine all historical events for this service
service_data = pd.concat([
df[df['service'] == service_name]
for df in historical_events
])
# Feature engineering
features = self._build_event_features(service_data, business_context)
targets = service_data['rps'].values
# Train ensemble model
model = RandomForestRegressor(
n_estimators=500,
max_depth=12,
min_samples_leaf=10,
n_jobs=-1,
)
model.fit(features, targets)
self.models[service_name] = model
# Calculate prediction intervals using OOB residuals
oob_predictions = model.oob_prediction_
residuals = targets - oob_predictions
metrics[service_name] = {
'mape': np.mean(np.abs(residuals / targets)) * 100,
'residual_std': np.std(residuals),
}
return metrics
def forecast_event(self,
event_date: pd.Timestamp,
business_signals: dict,
duration_hours: int = 24) -> list[CapacityForecast]:
"""Generate minute-by-minute capacity forecast for the event."""
forecasts = []
timestamps = pd.date_range(
start=event_date,
periods=duration_hours * 60,
freq='1min',
)
for ts in timestamps:
service_instances = {}
total_cost = 0.0
total_confidence = 0.0
for service_name, config in self.services.items():
model = self.models[service_name]
features = self._build_prediction_features(
ts, business_signals
)
# Point prediction
predicted_rps = model.predict(features.reshape(1, -1))[0]
# Prediction intervals via quantile estimation
tree_predictions = np.array([
tree.predict(features.reshape(1, -1))[0]
for tree in model.estimators_
])
p50 = np.percentile(tree_predictions, 50)
p75 = np.percentile(tree_predictions, 75)
p95 = np.percentile(tree_predictions, 95)
p99 = np.percentile(tree_predictions, 99)
# Convert RPS to instance count
capacity_per_instance = config['rps_capacity']
# Use P95 for provisioning (handles most scenarios)
instances_needed = int(np.ceil(
p95 / capacity_per_instance * 1.2 # 20% buffer
))
instances_needed = max(instances_needed, config['min_instances'])
service_instances[service_name] = instances_needed
total_cost += instances_needed * config['hourly_cost'] / 60
# Confidence = inverse of coefficient of variation
cv = np.std(tree_predictions) / np.mean(tree_predictions)
total_confidence += max(0, 1 - cv)
avg_confidence = total_confidence / len(self.services)
forecasts.append(CapacityForecast(
timestamp=ts,
predicted_rps=predicted_rps,
p50_rps=p50,
p75_rps=p75,
p95_rps=p95,
p99_rps=p99,
required_instances=service_instances,
estimated_cost=total_cost,
confidence=avg_confidence,
))
return forecasts
def _build_event_features(self, data, context) -> np.ndarray:
"""Build feature matrix from historical event data."""
ts = pd.DatetimeIndex(data['timestamp'])
features = pd.DataFrame({
'hour': ts.hour,
'minute': ts.minute,
'minutes_since_midnight': ts.hour * 60 + ts.minute,
'is_peak_window': ((ts.hour >= 8) & (ts.hour <= 22)).astype(int),
'hour_sin': np.sin(2 * np.pi * ts.hour / 24),
'hour_cos': np.cos(2 * np.pi * ts.hour / 24),
'year_growth': (ts.year - 2022).values, # YoY growth factor
})
return features.values
def _build_prediction_features(self, ts, signals) -> np.ndarray:
"""Build features for a single prediction timestamp."""
return np.array([
ts.hour,
ts.minute,
ts.hour * 60 + ts.minute,
1 if 8 <= ts.hour <= 22 else 0,
np.sin(2 * np.pi * ts.hour / 24),
np.cos(2 * np.pi * ts.hour / 24),
ts.year - 2022,
])
The Provisioning Plan Generator
The forecast becomes an actionable provisioning timeline — exactly when to scale each service and to what level:
@dataclass
class ProvisioningAction:
timestamp: pd.Timestamp
service: str
action: str # 'scale_up', 'scale_down', 'pre_warm'
target_instances: int
current_instances: int
lead_time_minutes: int
estimated_cost_delta: float
def generate_provisioning_plan(
forecasts: list[CapacityForecast],
current_state: dict,
boot_time_minutes: int = 5,
) -> list[ProvisioningAction]:
"""Convert forecasts into a time-ordered provisioning plan."""
actions = []
previous_targets = {svc: current_state[svc] for svc in current_state}
for forecast in forecasts:
for service, target in forecast.required_instances.items():
current = previous_targets.get(service, 0)
# Only generate action if change exceeds 10% threshold
if abs(target - current) / max(current, 1) < 0.1:
continue
if target > current:
# Scale up: schedule BEFORE the need (lead time)
action_time = forecast.timestamp - pd.Timedelta(
minutes=boot_time_minutes + 2 # Boot + health check
)
actions.append(ProvisioningAction(
timestamp=action_time,
service=service,
action='scale_up',
target_instances=target,
current_instances=current,
lead_time_minutes=boot_time_minutes,
estimated_cost_delta=(target - current) * 0.15, # $/hr per instance
))
elif target < current and current - target > 5:
# Scale down: only if significant reduction, with delay
action_time = forecast.timestamp + pd.Timedelta(minutes=15)
actions.append(ProvisioningAction(
timestamp=action_time,
service=service,
action='scale_down',
target_instances=target,
current_instances=current,
lead_time_minutes=0,
estimated_cost_delta=(target - current) * 0.15,
))
previous_targets[service] = target
# Deduplicate: keep only the last action per service per 10-min window
return _deduplicate_actions(actions)
Black Friday 2025: Actual vs Predicted
The model generated a provisioning timeline with 1,440 data points (one per minute for 24 hours). Here is how it performed:
| Service | Predicted Peak RPS | Actual Peak RPS | Provisioned Instances | Peak Utilization |
|---|---|---|---|---|
| API Gateway | 142,000 | 134,800 | 84 | 78% |
| Order Service | 28,000 | 31,200 | 42 | 86% |
| Payment Service | 18,000 | 16,400 | 24 | 72% |
| Search Service | 89,000 | 95,200 | 62 | 91% |
| Product Catalog | 210,000 | 198,000 | 48 | 82% |
The Search Service was the only near-miss: 91% utilization at peak, dangerously close to our 85% comfort threshold. The model underestimated search traffic because this year's promotional strategy featured a new "search and win" gamification that was not present in historical data.
Cost Comparison: AI-Driven vs Manual
| Approach | Total Event Cost | Peak Utilization | Over-Provisioned By |
|---|---|---|---|
| 2023 (manual, 80x) | $210K | 39% | 105% |
| 2024 (manual, 60x) | $180K | 63% | 58% |
| 2025 (AI-driven) | $52K | 82% avg | 18% |
The $128K savings came from two sources:
- Right-sized peak: Provisioned 47x instead of 60x ($82K saved)
- Shaped provisioning curve: Scaled up gradually and down after peak instead of flat-topping for 24 hours ($46K saved)
The Feedback Loop: Real-Time Adjustments
Even with a 30-day-ahead forecast, we run real-time adjustments on the day of the event. A lightweight model compares the morning's actual traffic shape to the predicted shape and adjusts afternoon forecasts:
def real_time_adjustment(
forecast: list[CapacityForecast],
actual_so_far: pd.DataFrame,
adjustment_timestamp: pd.Timestamp,
) -> list[CapacityForecast]:
"""Adjust remaining forecast based on observed vs predicted ratio."""
# Compare last 2 hours of actual vs predicted
lookback = adjustment_timestamp - pd.Timedelta(hours=2)
recent_actual = actual_so_far[actual_so_far['timestamp'] > lookback]
recent_forecast = [
f for f in forecast
if lookback < f.timestamp <= adjustment_timestamp
]
if not recent_forecast:
return forecast
# Calculate scaling factor
actual_mean = recent_actual['rps'].mean()
predicted_mean = np.mean([f.predicted_rps for f in recent_forecast])
if predicted_mean == 0:
return forecast
scale_factor = actual_mean / predicted_mean
# Apply dampened adjustment to future predictions
# Dampen to avoid overreacting to short-term variance
dampened_factor = 1.0 + (scale_factor - 1.0) * 0.6
adjusted = []
for f in forecast:
if f.timestamp > adjustment_timestamp:
adjusted.append(CapacityForecast(
timestamp=f.timestamp,
predicted_rps=f.predicted_rps * dampened_factor,
p50_rps=f.p50_rps * dampened_factor,
p75_rps=f.p75_rps * dampened_factor,
p95_rps=f.p95_rps * dampened_factor,
p99_rps=f.p99_rps * dampened_factor,
required_instances={
svc: int(np.ceil(count * dampened_factor))
for svc, count in f.required_instances.items()
},
estimated_cost=f.estimated_cost * dampened_factor,
confidence=f.confidence * 0.9, # Slight confidence reduction
))
else:
adjusted.append(f)
return adjusted
On Black Friday morning, the model detected that actual traffic was running 8% above forecast. The real-time adjustment bumped afternoon predictions up by ~5% (dampened), which triggered an additional 6 instances across services 45 minutes before the actual peak. Those 6 instances prevented what would have been a search service saturation event.
Lessons Learned
Three years of event data is the minimum. One year gives you a shape, two years give you a trend, three years give you confidence intervals. We could not build this system before 2025 because we only started collecting minute-level event data in 2022.
Business signals matter more than historical patterns. A 20% deeper promotional discount drives 35% more traffic. A new product category launch changes the traffic mix entirely. Feed your model the marketing plan, not just last year's metrics.
Test the provisioning plan with chaos engineering. Two weeks before Black Friday, we ran a simulated load test at 50x using the predicted traffic shape. This validated both the model's output and our infrastructure's ability to handle the predicted load.
Scale down conservatively. Our model aggressively scaled down after midnight, which was correct in 2024 but wrong in 2025 — international customers in different timezones created a longer tail. We added a 2-hour delay buffer for scale-down actions.
Conclusion
AI-assisted capacity planning transforms peak traffic events from existential threats into engineering problems with quantifiable solutions. The model is not magic — it is time-series forecasting with business context injection and real-time feedback. The value comes from replacing human intuition ("let's do 60x to be safe") with data-driven precision ("47x with shaped provisioning handles P95 scenarios"). For any organization spending more than $50K on peak-event infrastructure, the ML investment pays for itself in a single event cycle.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.