GitHub Actions Self-Hosted Runners on AWS with Auto-Scaling

Building a cost-efficient, auto-scaling GitHub Actions runner fleet on EC2 that reduced our CI/CD costs by 64% while cutting build times in half.

#github-actions#cicd#aws#automation
Cover image for the article: GitHub Actions Self-Hosted Runners on AWS with Auto-Scaling

GitHub-hosted runners are convenient until your bill crosses $15,000/month and your builds queue for 20 minutes during peak hours. We transitioned 180+ repositories to self-hosted runners on EC2 with auto-scaling and cut costs by 64% while reducing median build time from 14 minutes to 6.5 minutes. Here is the architecture that made it work.

The Problem: Cost and Performance at Scale

With 47 engineers pushing code across 180 repositories, our GitHub Actions usage looked like this:

  • Monthly spend: $15,800 on GitHub-hosted runners
  • Peak queue time: 22 minutes during morning standup deployments
  • Build cache misses: 73% due to ephemeral runner environments
  • GPU workloads: Impossible on standard GitHub runners, forcing a separate CI system

The fundamental limitation is that GitHub-hosted runners are general-purpose machines with no persistent state. Every job starts cold — downloading dependencies, building Docker layers, and compiling from scratch.

Architecture: Ephemeral Auto-Scaling Runner Fleet

Our architecture uses ephemeral EC2 instances managed by a Lambda-based scaler that responds to GitHub webhook events.

Self-Hosted Runner Architecture

Core Components

  1. Webhook Receiver: API Gateway + Lambda that processes workflow_job events
  2. Scaler: Lambda function that manages EC2 instance lifecycle
  3. Runner AMI: Pre-baked AMI with runners, Docker, and warm caches
  4. Instance Pool: Mix of on-demand and spot instances across multiple AZs

Webhook-Driven Scaling

The scaler responds to three webhook events: queued, in_progress, and completed.

// lambda/scaler/handler.ts
import { EC2Client, RunInstancesCommand, TerminateInstancesCommand } from '@aws-sdk/client-ec2';

interface WorkflowJobEvent {
  action: 'queued' | 'in_progress' | 'completed';
  workflow_job: {
    id: number;
    labels: string[];
    runner_group_name: string;
  };
}

export async function handler(event: WorkflowJobEvent): Promise<void> {
  const ec2 = new EC2Client({ region: process.env.AWS_REGION });
  const labels = event.workflow_job.labels;

  if (event.action === 'queued') {
    const instanceType = resolveInstanceType(labels);
    const ami = resolveAMI(labels);

    const command = new RunInstancesCommand({
      ImageId: ami,
      InstanceType: instanceType,
      MinCount: 1,
      MaxCount: 1,
      SubnetId: selectSubnet(labels),
      IamInstanceProfile: { Name: 'github-runner-profile' },
      InstanceMarketOptions: labels.includes('spot') ? {
        MarketType: 'spot',
        SpotOptions: {
          SpotInstanceType: 'one-time',
          InstanceInterruptionBehavior: 'terminate'
        }
      } : undefined,
      UserData: Buffer.from(generateUserData(event.workflow_job)).toString('base64'),
      TagSpecifications: [{
        ResourceType: 'instance',
        Tags: [
          { Key: 'Purpose', Value: 'github-runner' },
          { Key: 'JobId', Value: String(event.workflow_job.id) },
          { Key: 'Ephemeral', Value: 'true' }
        ]
      }]
    });

    await ec2.send(command);
  }

  if (event.action === 'completed') {
    await terminateRunnerInstance(ec2, event.workflow_job.id);
  }
}

function resolveInstanceType(labels: string[]): string {
  if (labels.includes('gpu')) return 'g5.xlarge';
  if (labels.includes('large')) return 'c6i.4xlarge';
  if (labels.includes('arm64')) return 'c7g.2xlarge';
  return 'c6i.2xlarge';
}

Pre-Baked Runner AMI

The AMI is built weekly with Packer and contains pre-warmed caches:

# packer/runner.pkr.hcl
source "amazon-ebs" "runner" {
  ami_name      = "github-runner-${formatdate("YYYYMMDD", timestamp())}"
  instance_type = "c6i.2xlarge"
  region        = "us-east-1"
  source_ami_filter {
    filters = {
      name                = "ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"
      virtualization-type = "hvm"
    }
    owners      = ["099720109477"]
    most_recent = true
  }
}

build {
  sources = ["source.amazon-ebs.runner"]

  provisioner "shell" {
    scripts = [
      "scripts/install-runner.sh",
      "scripts/install-docker.sh",
      "scripts/install-tools.sh",
      "scripts/warm-caches.sh"
    ]
  }

  provisioner "shell" {
    inline = [
      # Pre-pull common base images
      "docker pull node:20-alpine",
      "docker pull python:3.11-slim",
      "docker pull golang:1.22-alpine",
      "docker pull public.ecr.aws/lambda/provided:al2023",
      # Pre-install common npm packages into global cache
      "npm install -g typescript eslint prettier",
    ]
  }
}

Instance Lifecycle and Cost Optimization

Spot Instance Strategy

We run 80% of builds on spot instances with automatic fallback to on-demand:

Instance TypeUse CaseSpot SavingsInterruption Rate
c6i.2xlargeStandard builds67%3.2%
c7g.2xlargeARM builds71%1.8%
g5.xlargeGPU/ML workloads58%5.1%
c6i.4xlargeIntegration tests64%4.3%

When spot capacity is unavailable, the scaler falls back to on-demand with a 30-second timeout.

Warm Pool for Latency-Sensitive Workflows

Production deployment pipelines cannot wait for instance boot. We maintain a warm pool of 3 pre-registered runners:

# terraform/warm-pool.tf
resource "aws_autoscaling_group" "warm_runners" {
  name                = "github-runners-warm-pool"
  min_size            = 3
  max_size            = 3
  desired_capacity    = 3
  vpc_zone_identifier = var.private_subnet_ids

  launch_template {
    id      = aws_launch_template.runner.id
    version = "$Latest"
  }

  tag {
    key                 = "Pool"
    value               = "warm"
    propagate_at_launch = true
  }
}

Performance Benchmarks

Build Time Comparison

MetricGitHub-HostedSelf-HostedImprovement
Median build time14.2 min6.5 min54% faster
P95 build time28.4 min11.2 min61% faster
Peak queue wait22 min38 sec97% reduction
Docker layer cache hit27%94%3.5x improvement
Monthly cost$15,800$5,69064% savings
Concurrent capacity2060+3x increase

Operational Considerations

Security Isolation

Every runner is ephemeral — the instance is terminated immediately after job completion. No secrets persist between builds. Runners operate in private subnets with NAT Gateway egress only.

Monitoring and Alerting

We track four key signals:

  • Queue depth: Jobs waiting for a runner (target: < 2)
  • Boot-to-ready time: Instance launch to runner registration (target: < 90s)
  • Spot interruption rate: Weekly interruption percentage (alert: > 10%)
  • Cost per build minute: Normalized cost efficiency metric

Graceful Spot Interruption Handling

When AWS reclaims a spot instance, we catch the 2-minute warning and re-queue the job:

#!/bin/bash
# scripts/spot-interruption-handler.sh
METADATA_TOKEN=$(curl -X PUT "http://169.254.169.254/latest/api/token" \
  -H "X-aws-ec2-metadata-token-ttl-seconds: 30")

while true; do
  STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
    -H "X-aws-ec2-metadata-token: $METADATA_TOKEN" \
    http://169.254.169.254/latest/meta-data/spot/instance-action)

  if [ "$STATUS" -eq 200 ]; then
    echo "Spot interruption notice received. Draining runner..."
    /opt/actions-runner/bin/Runner.Listener remove --token "$RUNNER_TOKEN"
    # Signal the scaler to launch a replacement
    aws sns publish --topic-arn "$SCALER_TOPIC" \
      --message "{\"action\": \"spot-interrupted\", \"job_id\": \"$JOB_ID\"}"
    exit 0
  fi
  sleep 5
done

Key Takeaways

  1. Webhook-driven scaling beats polling: Responding to workflow_job.queued events gives you sub-minute scaling response versus 60-second polling intervals.

  2. Pre-bake everything possible: AMI boot time dominates cold-start latency. Every tool, runtime, and Docker image you pre-install saves seconds on every build.

  3. Spot instances are safe for CI: With proper interruption handling and on-demand fallback, spot provides massive savings with negligible reliability impact.

  4. Ephemeral runners simplify security: No persistent state means no credential leakage between builds and no drift in runner configuration.

  5. Measure cost per build minute: Raw EC2 cost is meaningless without normalization. Our cost per build minute dropped from $0.034 to $0.011 — a better efficiency metric than total spend.

The migration took 3 weeks of engineering time and paid for itself within the first billing cycle.

Comments

    No comments yet. Be the first to share your thoughts.