Shift-Left Container Security: Building a Zero-Vulnerability CI/CD Pipeline
How we implemented multi-layer container image scanning that blocked 340+ vulnerabilities from reaching production in 90 days

A critical CVE in a base image made it to production because our security scanning only ran weekly. By the time the scan flagged it, the vulnerable container had been serving traffic for nine days. We rebuilt our CI/CD pipeline to scan at every layer: Dockerfile lint, build-time dependency analysis, post-build image scan, and runtime admission control. In the first 90 days, the pipeline blocked 340+ vulnerabilities before they reached any environment beyond local development.
The Problem: Security as an Afterthought
Our previous approach was reactive. AWS ECR image scanning ran on push, but results were informational only. Nobody monitored the findings dashboard. Critical vulnerabilities accumulated silently until quarterly security reviews surfaced them, at which point remediation was expensive and disruptive.
Before shift-left scanning:
- 47 container images with known critical CVEs in production
- Average vulnerability age in production: 34 days
- No scanning during development or CI
- ECR scan results ignored by 89% of teams
- Zero policy enforcement on image security posture
Pipeline Architecture
The scanning pipeline operates at four checkpoints, each progressively more comprehensive:
Each layer catches different classes of issues. Dockerfile linting catches configuration weaknesses. Dependency scanning catches known library vulnerabilities. Image scanning catches OS-level CVEs. Admission control is the final gate preventing bypass.
Layer 1: Dockerfile Linting in Pre-Commit
The first line of defense catches insecure Dockerfile patterns before code leaves the developer's machine:
# .github/workflows/container-security.yml
name: Container Security Pipeline
on:
pull_request:
paths:
- '**/Dockerfile*'
- '**/docker-compose*.yml'
- '**/package-lock.json'
- '**/requirements*.txt'
- '**/go.sum'
jobs:
dockerfile-lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run Hadolint
uses: hadolint/hadolint-action@v3.1.0
with:
dockerfile: Dockerfile
failure-threshold: warning
ignore: DL3008 # Allow unpinned apt versions for base image updates
- name: Check for secrets in build args
run: |
if grep -rn "ARG.*PASSWORD\|ARG.*SECRET\|ARG.*TOKEN\|ARG.*KEY" Dockerfile*; then
echo "::error::Secrets detected in Dockerfile build arguments"
exit 1
fi
- name: Verify non-root user
run: |
if ! grep -q "^USER " Dockerfile; then
echo "::error::Dockerfile must specify a non-root USER"
exit 1
fi
This catches the most common container security mistakes: running as root, embedding secrets in build arguments, and using unverified base images.
Layer 2: Dependency Scanning with Trivy
After Dockerfile validation, we scan application dependencies for known vulnerabilities:
dependency-scan:
runs-on: ubuntu-latest
needs: dockerfile-lint
steps:
- uses: actions/checkout@v4
- name: Run Trivy filesystem scan
uses: aquasecurity/trivy-action@master
with:
scan-type: 'fs'
scan-ref: '.'
severity: 'CRITICAL,HIGH'
exit-code: '1'
format: 'sarif'
output: 'trivy-fs-results.sarif'
- name: Upload SARIF results
uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: 'trivy-fs-results.sarif'
Layer 3: Built Image Scanning
The most comprehensive scan runs against the fully built container image. We use both Trivy and Grype for defense-in-depth, since each scanner has different vulnerability database coverage:
import subprocess
import json
import sys
from dataclasses import dataclass
@dataclass
class ScanResult:
scanner: str
critical: int
high: int
medium: int
low: int
fixable_critical: int
fixable_high: int
def scan_image(image_ref: str) -> list[ScanResult]:
"""Run multiple scanners against a container image."""
results = []
# Trivy scan
trivy_output = subprocess.run(
[
"trivy", "image",
"--format", "json",
"--severity", "CRITICAL,HIGH,MEDIUM,LOW",
"--vuln-type", "os,library",
image_ref,
],
capture_output=True, text=True
)
trivy_data = json.loads(trivy_output.stdout)
results.append(parse_trivy_results(trivy_data))
# Grype scan
grype_output = subprocess.run(
["grype", image_ref, "-o", "json"],
capture_output=True, text=True
)
grype_data = json.loads(grype_output.stdout)
results.append(parse_grype_results(grype_data))
return results
def enforce_policy(results: list[ScanResult]) -> bool:
"""Enforce security policy against scan results.
Policy:
- Zero critical vulnerabilities with available fixes
- Maximum 5 high vulnerabilities with available fixes
- All unfixable vulnerabilities must have accepted risk tickets
"""
for result in results:
if result.fixable_critical > 0:
print(f"BLOCKED: {result.scanner} found {result.fixable_critical} "
f"fixable critical vulnerabilities")
return False
if result.fixable_high > 5:
print(f"BLOCKED: {result.scanner} found {result.fixable_high} "
f"fixable high vulnerabilities (max: 5)")
return False
return True
if __name__ == "__main__":
image = sys.argv[1]
results = scan_image(image)
if not enforce_policy(results):
sys.exit(1)
print("PASSED: Image meets security policy requirements")
The key policy decision: we only block on fixable vulnerabilities. Unfixable CVEs (no patch available) follow a risk acceptance workflow rather than blocking deployments indefinitely.
Layer 4: Kubernetes Admission Control
Even with CI/CD scanning, images can bypass the pipeline (emergency deploys, manual kubectl applies). The admission controller is the final enforcement point:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
name: image-security-webhook
webhooks:
- name: image-policy.security.company.com
rules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
- apiGroups: ["apps"]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["deployments", "statefulsets", "daemonsets"]
clientConfig:
service:
name: image-policy-webhook
namespace: security-system
path: /validate
failurePolicy: Fail
sideEffects: None
admissionReviewVersions: ["v1"]
---
apiVersion: v1
kind: ConfigMap
metadata:
name: image-policy-config
namespace: security-system
data:
policy.yaml: |
rules:
- name: require-scan-attestation
description: "All images must have a passing scan attestation"
match:
namespaces: ["production", "staging"]
require:
attestations:
- predicateType: "https://cosign.sigstore.dev/attestation/vuln/v1"
conditions:
- all:
- key: "{{ .scanner }}"
operator: In
values: ["trivy", "grype"]
- key: "{{ .fixable_critical }}"
operator: Equals
value: "0"
- name: require-image-signature
description: "All images must be signed"
match:
namespaces: ["production"]
require:
signatures:
- keyless:
issuer: "https://accounts.google.com"
subject: "ci-pipeline@company.iam.gserviceaccount.com"
The admission controller validates two things: the image has a vulnerability scan attestation showing zero fixable criticals, and the image is cryptographically signed by our CI pipeline. Manual pushes to ECR without going through CI are rejected at deploy time.
Base Image Management
We maintain a curated set of approved base images that receive automated patching:
A weekly automated pipeline rebuilds all approved base images with the latest security patches, scans them, and publishes new tags. Application teams pin to our internal base images rather than pulling directly from Docker Hub.
Results: 90-Day Metrics
| Metric | Before | After | Change |
|---|---|---|---|
| Critical CVEs in production | 47 | 0 | -100% |
| Average vulnerability age | 34 days | 0 days (blocked) | -100% |
| Time to patch critical CVE | 11 days | 4.2 hours | -98% |
| Vulnerabilities blocked in CI | 0 | 340+ | N/A |
| False positive rate | N/A | 3.2% | Acceptable |
| Pipeline time added | N/A | +2.8 minutes | Acceptable |
Managing False Positives
A 3.2% false positive rate means roughly 11 of our 340 blocks were unnecessary. We manage this with a .trivyignore file and a risk acceptance process:
# Accepted risks with justification and expiry
# Format: CVE-ID # Reason | Owner | Expires
CVE-2024-1234 # No fix available, mitigated by WAF rules | @security-team | 2026-07-01
CVE-2024-5678 # Disputed, not exploitable in our configuration | @platform-team | 2026-06-15
Accepted risks expire automatically. If the expiry passes without renewal, the vulnerability blocks deployments again, forcing periodic re-evaluation.
Lessons Learned
Block on fixable vulnerabilities only. Blocking on all vulnerabilities creates alert fatigue and incentivizes workarounds. Block what can be fixed, track what cannot.
Dual-scanner coverage catches more. Trivy and Grype have different vulnerability databases. Running both caught 12% more issues than either alone.
Admission control is non-negotiable. CI scanning alone had a 7% bypass rate from emergency deploys and manual interventions. The admission controller closed that gap completely.
Base image management is upstream leverage. Fixing a vulnerability in a base image automatically fixes it in 23 downstream application images. Invest heavily in base image hygiene.
Conclusion
Shift-left container security is not about moving a single scan earlier. It is about layered defense: catch what you can early, enforce policy at build time, and maintain a final gate at runtime. The 340 blocked vulnerabilities in 90 days proved that our previous approach was not just slow but blind. Zero fixable critical CVEs in production is achievable with the right pipeline architecture.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.