GitHub Actions Self-Hosted Runners on AWS with Auto-Scaling
Building a cost-efficient, auto-scaling GitHub Actions runner fleet on EC2 that reduced our CI/CD costs by 64% while cutting build times in half.

GitHub-hosted runners are convenient until your bill crosses $15,000/month and your builds queue for 20 minutes during peak hours. We transitioned 180+ repositories to self-hosted runners on EC2 with auto-scaling and cut costs by 64% while reducing median build time from 14 minutes to 6.5 minutes. Here is the architecture that made it work.
The Problem: Cost and Performance at Scale
With 47 engineers pushing code across 180 repositories, our GitHub Actions usage looked like this:
- Monthly spend: $15,800 on GitHub-hosted runners
- Peak queue time: 22 minutes during morning standup deployments
- Build cache misses: 73% due to ephemeral runner environments
- GPU workloads: Impossible on standard GitHub runners, forcing a separate CI system
The fundamental limitation is that GitHub-hosted runners are general-purpose machines with no persistent state. Every job starts cold — downloading dependencies, building Docker layers, and compiling from scratch.
Architecture: Ephemeral Auto-Scaling Runner Fleet
Our architecture uses ephemeral EC2 instances managed by a Lambda-based scaler that responds to GitHub webhook events.
Core Components
- Webhook Receiver: API Gateway + Lambda that processes
workflow_jobevents - Scaler: Lambda function that manages EC2 instance lifecycle
- Runner AMI: Pre-baked AMI with runners, Docker, and warm caches
- Instance Pool: Mix of on-demand and spot instances across multiple AZs
Webhook-Driven Scaling
The scaler responds to three webhook events: queued, in_progress, and completed.
// lambda/scaler/handler.ts
import { EC2Client, RunInstancesCommand, TerminateInstancesCommand } from '@aws-sdk/client-ec2';
interface WorkflowJobEvent {
action: 'queued' | 'in_progress' | 'completed';
workflow_job: {
id: number;
labels: string[];
runner_group_name: string;
};
}
export async function handler(event: WorkflowJobEvent): Promise<void> {
const ec2 = new EC2Client({ region: process.env.AWS_REGION });
const labels = event.workflow_job.labels;
if (event.action === 'queued') {
const instanceType = resolveInstanceType(labels);
const ami = resolveAMI(labels);
const command = new RunInstancesCommand({
ImageId: ami,
InstanceType: instanceType,
MinCount: 1,
MaxCount: 1,
SubnetId: selectSubnet(labels),
IamInstanceProfile: { Name: 'github-runner-profile' },
InstanceMarketOptions: labels.includes('spot') ? {
MarketType: 'spot',
SpotOptions: {
SpotInstanceType: 'one-time',
InstanceInterruptionBehavior: 'terminate'
}
} : undefined,
UserData: Buffer.from(generateUserData(event.workflow_job)).toString('base64'),
TagSpecifications: [{
ResourceType: 'instance',
Tags: [
{ Key: 'Purpose', Value: 'github-runner' },
{ Key: 'JobId', Value: String(event.workflow_job.id) },
{ Key: 'Ephemeral', Value: 'true' }
]
}]
});
await ec2.send(command);
}
if (event.action === 'completed') {
await terminateRunnerInstance(ec2, event.workflow_job.id);
}
}
function resolveInstanceType(labels: string[]): string {
if (labels.includes('gpu')) return 'g5.xlarge';
if (labels.includes('large')) return 'c6i.4xlarge';
if (labels.includes('arm64')) return 'c7g.2xlarge';
return 'c6i.2xlarge';
}
Pre-Baked Runner AMI
The AMI is built weekly with Packer and contains pre-warmed caches:
# packer/runner.pkr.hcl
source "amazon-ebs" "runner" {
ami_name = "github-runner-${formatdate("YYYYMMDD", timestamp())}"
instance_type = "c6i.2xlarge"
region = "us-east-1"
source_ami_filter {
filters = {
name = "ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"
virtualization-type = "hvm"
}
owners = ["099720109477"]
most_recent = true
}
}
build {
sources = ["source.amazon-ebs.runner"]
provisioner "shell" {
scripts = [
"scripts/install-runner.sh",
"scripts/install-docker.sh",
"scripts/install-tools.sh",
"scripts/warm-caches.sh"
]
}
provisioner "shell" {
inline = [
# Pre-pull common base images
"docker pull node:20-alpine",
"docker pull python:3.11-slim",
"docker pull golang:1.22-alpine",
"docker pull public.ecr.aws/lambda/provided:al2023",
# Pre-install common npm packages into global cache
"npm install -g typescript eslint prettier",
]
}
}
Instance Lifecycle and Cost Optimization
Spot Instance Strategy
We run 80% of builds on spot instances with automatic fallback to on-demand:
| Instance Type | Use Case | Spot Savings | Interruption Rate |
|---|---|---|---|
| c6i.2xlarge | Standard builds | 67% | 3.2% |
| c7g.2xlarge | ARM builds | 71% | 1.8% |
| g5.xlarge | GPU/ML workloads | 58% | 5.1% |
| c6i.4xlarge | Integration tests | 64% | 4.3% |
When spot capacity is unavailable, the scaler falls back to on-demand with a 30-second timeout.
Warm Pool for Latency-Sensitive Workflows
Production deployment pipelines cannot wait for instance boot. We maintain a warm pool of 3 pre-registered runners:
# terraform/warm-pool.tf
resource "aws_autoscaling_group" "warm_runners" {
name = "github-runners-warm-pool"
min_size = 3
max_size = 3
desired_capacity = 3
vpc_zone_identifier = var.private_subnet_ids
launch_template {
id = aws_launch_template.runner.id
version = "$Latest"
}
tag {
key = "Pool"
value = "warm"
propagate_at_launch = true
}
}
Performance Benchmarks
| Metric | GitHub-Hosted | Self-Hosted | Improvement |
|---|---|---|---|
| Median build time | 14.2 min | 6.5 min | 54% faster |
| P95 build time | 28.4 min | 11.2 min | 61% faster |
| Peak queue wait | 22 min | 38 sec | 97% reduction |
| Docker layer cache hit | 27% | 94% | 3.5x improvement |
| Monthly cost | $15,800 | $5,690 | 64% savings |
| Concurrent capacity | 20 | 60+ | 3x increase |
Operational Considerations
Security Isolation
Every runner is ephemeral — the instance is terminated immediately after job completion. No secrets persist between builds. Runners operate in private subnets with NAT Gateway egress only.
Monitoring and Alerting
We track four key signals:
- Queue depth: Jobs waiting for a runner (target: < 2)
- Boot-to-ready time: Instance launch to runner registration (target: < 90s)
- Spot interruption rate: Weekly interruption percentage (alert: > 10%)
- Cost per build minute: Normalized cost efficiency metric
Graceful Spot Interruption Handling
When AWS reclaims a spot instance, we catch the 2-minute warning and re-queue the job:
#!/bin/bash
# scripts/spot-interruption-handler.sh
METADATA_TOKEN=$(curl -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 30")
while true; do
STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
-H "X-aws-ec2-metadata-token: $METADATA_TOKEN" \
http://169.254.169.254/latest/meta-data/spot/instance-action)
if [ "$STATUS" -eq 200 ]; then
echo "Spot interruption notice received. Draining runner..."
/opt/actions-runner/bin/Runner.Listener remove --token "$RUNNER_TOKEN"
# Signal the scaler to launch a replacement
aws sns publish --topic-arn "$SCALER_TOPIC" \
--message "{\"action\": \"spot-interrupted\", \"job_id\": \"$JOB_ID\"}"
exit 0
fi
sleep 5
done
Key Takeaways
-
Webhook-driven scaling beats polling: Responding to
workflow_job.queuedevents gives you sub-minute scaling response versus 60-second polling intervals. -
Pre-bake everything possible: AMI boot time dominates cold-start latency. Every tool, runtime, and Docker image you pre-install saves seconds on every build.
-
Spot instances are safe for CI: With proper interruption handling and on-demand fallback, spot provides massive savings with negligible reliability impact.
-
Ephemeral runners simplify security: No persistent state means no credential leakage between builds and no drift in runner configuration.
-
Measure cost per build minute: Raw EC2 cost is meaningless without normalization. Our cost per build minute dropped from $0.034 to $0.011 — a better efficiency metric than total spend.
The migration took 3 weeks of engineering time and paid for itself within the first billing cycle.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.