GCP Cloud Build CI/CD Pipeline Optimization
Optimizing Google Cloud Build pipelines for faster builds, reduced costs, and reliable deployments with caching, parallelism, and artifact management

Introduction
GCP Cloud Build is a serverless CI/CD platform that executes build steps as containers. While its simplicity is appealing, naive configurations result in slow builds, redundant work, and inflated costs. Production pipelines processing 200+ builds per day can see build times reduced by 60-75% through proper caching, step parallelism, and machine type selection.
This article presents optimization techniques measured against real pipeline metrics, with before-and-after comparisons for common build patterns.
Build Performance Baseline
Before optimization, a typical microservice pipeline:
| Stage | Duration (Before) | Duration (After) | Improvement |
|---|---|---|---|
| Source checkout | 15s | 5s | 67% |
| Dependency install | 120s | 20s | 83% |
| Unit tests | 90s | 45s | 50% |
| Build container | 180s | 45s | 75% |
| Push to registry | 30s | 10s | 67% |
| Deploy to staging | 60s | 60s | 0% |
| Integration tests | 120s | 120s | 0% |
| Total | 615s (10.25 min) | 305s (5.08 min) | 50% |
Machine Type Selection
Cloud Build offers multiple machine types. Selecting the right one depends on workload characteristics:
| Machine Type | vCPUs | Memory | Disk | Cost/min | Best For |
|---|---|---|---|---|---|
| E2_MEDIUM | 1 | 4 GB | 100 GB | $0.003 | Simple builds |
| E2_HIGHCPU_8 | 8 | 8 GB | 100 GB | $0.016 | Parallel compilation |
| E2_HIGHCPU_32 | 32 | 32 GB | 100 GB | $0.064 | Large monorepos |
| N1_HIGHCPU_8 | 8 | 7.2 GB | 200 GB | $0.016 | Standard pipelines |
| N1_HIGHCPU_32 | 32 | 28.8 GB | 400 GB | $0.064 | Heavy compilation |
Cost-Time Tradeoff
# Faster machine, lower total cost due to reduced build minutes
options:
machineType: 'E2_HIGHCPU_8'
diskSizeGb: 200
| Scenario | E2_MEDIUM (10 min) | E2_HIGHCPU_8 (3 min) | Savings |
|---|---|---|---|
| Per build | $0.03 | $0.048 | -60% (costs more) |
| With concurrency (200 builds/day) | $6.00/day | $9.60/day | -60% |
| Developer wait time saved | - | 7 min/build | 1,400 min/day |
The faster machine costs more per build but the developer productivity gain (1,400 minutes/day saved across team) far exceeds the $3.60/day additional cost.
Docker Layer Caching
The most impactful optimization for container builds. Use kaniko with cache enabled:
steps:
# Build with kaniko (supports layer caching)
- name: 'gcr.io/kaniko-project/executor:latest'
args:
- '--destination=${_REGION}-docker.pkg.dev/${PROJECT_ID}/${_REPO}/${_IMAGE}:${SHORT_SHA}'
- '--cache=true'
- '--cache-ttl=168h'
- '--cache-repo=${_REGION}-docker.pkg.dev/${PROJECT_ID}/${_REPO}/cache'
- '--context=.'
- '--dockerfile=Dockerfile'
- '--build-arg=VERSION=${SHORT_SHA}'
- '--snapshotMode=redo'
- '--compressed-caching=false'
Dockerfile Optimization for Caching
# Stage 1: Dependencies (cached unless package.json changes)
FROM node:20-alpine AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --production=false
# Stage 2: Build (cached unless source changes)
FROM deps AS builder
COPY tsconfig.json ./
COPY src/ ./src/
RUN npm run build
# Stage 3: Production image (minimal)
FROM node:20-alpine AS runner
WORKDIR /app
RUN addgroup --system --gid 1001 nodejs && \
adduser --system --uid 1001 appuser
COPY --from=builder /app/dist ./dist
COPY --from=deps /app/node_modules ./node_modules
COPY package.json ./
USER appuser
EXPOSE 8080
CMD ["node", "dist/index.js"]
Cache Hit Rate Impact
| Scenario | Build Time (no cache) | Build Time (cached) | Improvement |
|---|---|---|---|
| No code changes | 180s | 15s | 92% |
| Source code change only | 180s | 45s | 75% |
| Dependency change | 180s | 90s | 50% |
| Dockerfile change | 180s | 180s | 0% |
Parallel Step Execution
Cloud Build supports waitFor to execute steps in parallel:
steps:
# Step 0: Checkout (all steps wait for this)
- id: 'checkout'
name: 'gcr.io/cloud-builders/git'
args: ['clone', '--depth=1', '${_REPO_URL}', '.']
# Steps 1-3 run in PARALLEL (all wait only for checkout)
- id: 'lint'
name: 'node:20-alpine'
entrypoint: 'sh'
args: ['-c', 'npm ci && npm run lint']
waitFor: ['checkout']
- id: 'unit-tests'
name: 'node:20-alpine'
entrypoint: 'sh'
args: ['-c', 'npm ci && npm run test:unit']
waitFor: ['checkout']
- id: 'security-scan'
name: 'snyk/snyk:node'
args: ['test', '--severity-threshold=high']
waitFor: ['checkout']
# Step 4: Build (waits for lint + tests to pass)
- id: 'build'
name: 'gcr.io/kaniko-project/executor:latest'
args:
- '--destination=${_REGION}-docker.pkg.dev/${PROJECT_ID}/${_REPO}/${_IMAGE}:${SHORT_SHA}'
- '--cache=true'
waitFor: ['lint', 'unit-tests', 'security-scan']
# Step 5: Deploy (waits for build)
- id: 'deploy'
name: 'gcr.io/google.com/cloudsdktool/cloud-sdk'
entrypoint: 'gcloud'
args: ['run', 'deploy', '${_SERVICE}', '--image=${_IMAGE}:${SHORT_SHA}', '--region=${_REGION}']
waitFor: ['build']
options:
machineType: 'E2_HIGHCPU_8'
substitutions:
_REGION: us-central1
_REPO: docker-repo
_IMAGE: my-service
_SERVICE: my-service
Dependency Caching with GCS
For package managers that do not benefit from Docker layer caching:
steps:
# Restore cache
- id: 'restore-cache'
name: 'gcr.io/cloud-builders/gsutil'
args: ['cp', 'gs://${PROJECT_ID}-build-cache/node_modules.tar.gz', '/workspace/']
allowFailure: true
- id: 'extract-cache'
name: 'ubuntu'
entrypoint: 'bash'
args:
- '-c'
- |
if [ -f /workspace/node_modules.tar.gz ]; then
tar xzf /workspace/node_modules.tar.gz
echo "Cache restored"
else
echo "No cache found"
fi
waitFor: ['restore-cache']
# Install dependencies (fast if cached)
- id: 'install'
name: 'node:20-alpine'
entrypoint: 'sh'
args: ['-c', 'npm ci']
waitFor: ['extract-cache']
# Save cache (only if lock file changed)
- id: 'save-cache'
name: 'gcr.io/cloud-builders/gsutil'
entrypoint: 'bash'
args:
- '-c'
- |
tar czf /workspace/node_modules.tar.gz node_modules/
gsutil cp /workspace/node_modules.tar.gz gs://${PROJECT_ID}-build-cache/
waitFor: ['install']
Build Triggers and Branch Strategy
# cloudbuild-trigger.yaml
trigger:
name: "main-deploy"
github:
owner: "my-org"
name: "my-repo"
push:
branch: "^main$"
includedFiles:
- "src/**"
- "package.json"
- "Dockerfile"
ignoredFiles:
- "docs/**"
- "*.md"
- ".github/**"
Trigger Strategy by Branch
| Branch Pattern | Trigger Action | Machine Type | Timeout |
|---|---|---|---|
| main | Build + Deploy (prod) | E2_HIGHCPU_8 | 30 min |
| release/* | Build + Deploy (staging) | E2_HIGHCPU_8 | 20 min |
| feature/* | Lint + Test only | E2_MEDIUM | 10 min |
| PR to main | Full pipeline (no deploy) | E2_HIGHCPU_8 | 20 min |
Cost Optimization Summary
| Optimization | Monthly Savings (200 builds/day) | Implementation Effort |
|---|---|---|
| Docker layer caching | $180/mo (reduced build minutes) | Low |
| Parallel steps | $120/mo (reduced wall time) | Low |
| Machine type right-sizing | $90/mo | Low |
| Trigger file filters | $60/mo (fewer unnecessary builds) | Low |
| GCS dependency cache | $45/mo | Medium |
| Total | ~$495/mo |
Key Takeaways
- Docker layer caching with kaniko provides 50-92% build time reduction depending on what changed between builds, making it the highest-impact single optimization.
- Use waitFor for parallel step execution to run lint, tests, and security scans simultaneously rather than sequentially.
- Select E2_HIGHCPU_8 for most pipelines as the faster build times offset higher per-minute costs through reduced developer wait time.
- Optimize Dockerfiles for cache hit rates by separating dependency installation from source code copying in multi-stage builds.
- Use trigger file filters to skip builds when only documentation or unrelated files change.
- Total optimization reduces build times by 50%+ and monthly costs by $400-500 for teams running 200+ builds per day.
- Monitor cache hit rates as a leading indicator of pipeline efficiency and investigate drops that indicate Dockerfile or dependency changes.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.