Docker Multi-Stage Build Optimization: Reducing Image Sizes by 78%
A systematic approach to multi-stage Docker builds that reduced our production image sizes from 1.2GB to 267MB while improving build cache efficiency.

Container image size directly impacts deployment speed, cold start latency, and infrastructure cost. When your Kubernetes cluster pulls a 1.2GB image across 200 pods during a rolling update, you are burning time, bandwidth, and money. We systematically optimized our Docker builds across 85 microservices, achieving a fleet-wide average reduction of 78%. Here is the methodology.
The Problem: Image Bloat at Fleet Scale
Our production fleet ran 85 containerized microservices. A routine audit revealed alarming numbers:
- Average image size: 1.2GB
- Largest image: 3.4GB (ML inference service)
- Pull time on cold nodes: 47 seconds average
- ECR storage cost: $2,800/month for image versions
- Rolling update duration: 8.5 minutes (dominated by image pull)
The root cause was consistent: developers were building images from full development environments. Build tools, test frameworks, source code, and dev dependencies all shipped to production.
Architecture: The Multi-Stage Build Pipeline
Multi-stage builds solve this by separating the build environment from the runtime environment. But naive multi-stage builds still leave significant optimization on the table.
The Four-Stage Pattern
We standardized on a four-stage pattern across all services:
# Stage 1: Dependencies (cached aggressively)
FROM node:20-alpine AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --only=production && \
cp -R node_modules /prod_modules && \
npm ci
# Stage 2: Build (compile TypeScript, bundle assets)
FROM node:20-alpine AS build
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY . .
RUN npm run build && \
npm prune --production
# Stage 3: Security scan (gate on vulnerabilities)
FROM aquasec/trivy:latest AS scan
COPY --from=build /app /scan-target
RUN trivy filesystem --exit-code 1 --severity HIGH,CRITICAL \
--no-progress /scan-target
# Stage 4: Production runtime (minimal)
FROM gcr.io/distroless/nodejs20-debian12 AS production
WORKDIR /app
COPY --from=deps /prod_modules ./node_modules
COPY --from=build /app/dist ./dist
COPY --from=build /app/package.json ./
USER nonroot:nonroot
EXPOSE 8080
CMD ["dist/server.js"]
Why Distroless Over Alpine
We migrated from Alpine to Distroless for production stages:
| Metric | Alpine-based | Distroless | Improvement |
|---|---|---|---|
| Base image size | 78MB | 43MB | 45% smaller |
| CVE surface | 12 packages | 0 shell/pkg manager | Minimal |
| Attack surface | Shell, apk, busybox | No shell | Significantly reduced |
| Debug capability | Full shell | Debug variant only | Controlled access |
The absence of a shell in production images eliminates an entire class of container escape vulnerabilities.
Layer Optimization Techniques
Dependency Layer Caching
The single most impactful optimization is isolating dependency installation into its own layer. When only source code changes, Docker reuses the cached dependency layer.
# Bad: Cache invalidated on ANY file change
COPY . .
RUN npm ci
# Good: Cache invalidated only when lock file changes
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
This reduced our average build time from 4.2 minutes to 1.1 minutes for code-only changes.
Binary Stripping and Compression
For Go services, we apply aggressive binary optimization:
FROM golang:1.22-alpine AS build
WORKDIR /app
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 GOOS=linux GOARCH=amd64 \
go build -ldflags="-s -w -X main.version=${VERSION}" \
-o /app/server ./cmd/server
# UPX compression for additional 60% reduction
RUN apk add --no-cache upx && \
upx --best --lzma /app/server
FROM scratch AS production
COPY --from=build /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/
COPY --from=build /app/server /server
USER 65534:65534
ENTRYPOINT ["/server"]
A typical Go binary drops from 45MB to 8MB after stripping and UPX compression.
Python Service Optimization
Python services present unique challenges due to compiled C extensions:
FROM python:3.11-slim AS build
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc libpq-dev && \
rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install --user --no-cache-dir -r requirements.txt
FROM python:3.11-slim AS production
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
libpq5 && \
rm -rf /var/lib/apt/lists/* && \
useradd --create-home --shell /bin/false appuser
COPY --from=build /root/.local /home/appuser/.local
COPY . .
USER appuser
ENV PATH=/home/appuser/.local/bin:$PATH
CMD ["gunicorn", "app:create_app()", "--bind", "0.0.0.0:8080"]
Key insight: install build dependencies (gcc, dev headers) only in the build stage. Copy only the compiled wheels to production.
Fleet-Wide Results
After rolling out optimized builds across 85 services over 6 weeks:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Average image size | 1.2GB | 267MB | 78% reduction |
| Largest image (ML) | 3.4GB | 890MB | 74% reduction |
| Cold node pull time | 47s | 11s | 77% faster |
| Rolling update duration | 8.5 min | 2.1 min | 75% faster |
| ECR monthly storage | $2,800 | $620 | 78% savings |
| Build cache hit rate | 34% | 87% | 2.6x improvement |
| Average build time | 4.2 min | 1.4 min | 67% faster |
CI Integration: Automated Size Enforcement
We enforce image size budgets in CI to prevent regression:
# .github/workflows/docker-build.yml
- name: Build and check image size
run: |
docker build -t $IMAGE_TAG .
SIZE=$(docker image inspect $IMAGE_TAG --format='{{.Size}}')
MAX_SIZE=$((500 * 1024 * 1024)) # 500MB budget
if [ "$SIZE" -gt "$MAX_SIZE" ]; then
echo "::error::Image size $(numfmt --to=iec $SIZE) exceeds budget of 500MB"
docker history $IMAGE_TAG --no-trunc
exit 1
fi
echo "Image size: $(numfmt --to=iec $SIZE) - within budget"
BuildKit Cache Mounts
Docker BuildKit cache mounts persist package manager caches across builds without including them in the final image:
# syntax=docker/dockerfile:1
FROM node:20-alpine AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm \
npm ci --prefer-offline
This eliminates redundant downloads during iterative development while keeping the layer clean.
Common Anti-Patterns We Eliminated
- Installing dev dependencies in production: Separating
npm ci --only=productionfor runtime dependencies. - Leaving build artifacts: Source maps, test files, and documentation that ship unnecessarily.
- Using full OS base images:
ubuntu:22.04(78MB) whenscratchor distroless (0-43MB) suffices. - Ignoring .dockerignore: Copying
.git/,node_modules/, and test fixtures into the build context. - Single-layer installs: Combining unrelated operations that break cache efficiency.
Key Takeaways
-
Measure before optimizing: Run
docker historyanddiveon your images to identify the largest layers. Optimize the biggest contributors first. -
Standardize a multi-stage template: Create team-wide Dockerfile templates per language. Consistency enables automated optimization and auditing.
-
Enforce budgets in CI: Image size regression is silent and cumulative. Automated gates catch bloat before it ships.
-
Cache strategically: Layer ordering determines cache efficiency. Put rarely-changing layers (OS, system deps) first, frequently-changing layers (source code) last.
-
Choose minimal base images: Distroless or scratch for production. The convenience of a shell is not worth the security and size cost.
The 78% reduction translated directly to faster deployments, lower infrastructure costs, and reduced attack surface — three wins from one systematic optimization pass.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.