Skip to main content

CI/CD for Docker Images: Automate Build, Testing, and Deployment

Calculating read time…

CI/CD for Docker images is the automated pipeline that turns a Dockerfile change into a scanned, tagged, attested, and pushed image ready for deployment — and doing it well means treating the image itself, not just the application code inside it, as the artifact that needs testing, versioning, and security gates. 🏗️

Teams that bolt Docker onto an existing CI/CD pipeline as an afterthought tend to hit the same wall: builds that take twenty minutes because nothing is cached, a latest tag that makes rollback guesswork, and a vulnerability discovered in production that a five-second scan would have caught before it ever shipped. None of this is exotic — it's a small, well-understood set of pipeline stages, done in the right order, with the right gates. Get the order wrong, and you either slow every developer down or ship risk straight into production. 🚦

Note: this post focuses on pipeline mechanics and doesn't include a diagram, consistent with how the last few posts in this series were scoped — happy to add one if it would help.

🔀 Quick Comparison: Pipeline Stages and What They're For

Each stage answers a different question about the image before it's trusted enough to deploy.

Stage Question it answers Typical tooling
Build Does the image build reproducibly from this commit? BuildKit, Buildx, multi-stage Dockerfiles
Tag How will this exact image be referenced later? Git SHA, semantic version, metadata-generation tooling
Scan Does this image contain known, unacceptable vulnerabilities? Docker Scout, other CVE scanners
Attest What's inside this image, and how was it built? SBOM (SPDX/CycloneDX), SLSA provenance
Push Where does the deployable artifact live? Docker Hub, GHCR, private registry

🎯 Use this table as a checklist when reviewing whether an existing pipeline is missing a stage, not just when building one from scratch.

1. Foundations — The Inner Loop, the Outer Loop, and Why Images Need Their Own Pipeline Discipline

🧸 Kid-friendly analogy: Think of baking cookies at home to taste-test a recipe (your inner loop) versus baking the same recipe in a commercial kitchen to sell to customers (your outer loop, the CI). If the commercial kitchen uses different ovens, different measuring cups, or a different recipe card than the one you tested at home, you'll be surprised by the results — and not in a good way.

Docker's own guidance frames developer workflow in terms of an inner loop — the local cycle of coding, building, running, and testing — and an outer loop, the CI process that builds, tests, and deploys once code is pushed. The core principle: the closer your inner loop resembles your outer loop, the fewer surprises you get in CI, because you're not "debugging through the CI" for problems your local build would have shown you immediately.

CI/CD for Docker images deserves its own discipline, distinct from ordinary application CI, for a few concrete reasons:

  • The image is the deployable artifact. Unlike a compiled binary or a set of test results, the container image itself — with its exact base layers, dependencies, and configuration — is what actually runs in production, so it needs to be versioned and audited as a first-class artifact.
  • Build reproducibility is a security property, not just a convenience. An image built from a floating base tag or unpinned dependencies can differ between builds of the identical source code, which undermines the ability to reason about what's actually running.
  • Scanning and attestation are part of "done," not an optional add-on. A build that produces a working image but skips vulnerability scanning and provenance recording has only completed half the job a production-grade pipeline needs to do.

2. Pipeline Mechanics — Build, Tag, Scan, Attest, Push

🧸 Kid-friendly analogy: Shipping a package internationally isn't just "put it in a box." You label it clearly (tag), have it inspected (scan), attach a customs declaration listing exactly what's inside (attest), and only then hand it to the shipping carrier (push). Skip any step and the package either gets lost, gets stuck, or causes a problem you only discover after it's already crossed the border.

What it does: a well-formed Docker CI/CD pipeline runs a fixed sequence of stages every time a change is pushed: build the image with BuildKit/Buildx, generate deterministic tags and metadata, scan for known vulnerabilities, attach supply-chain attestations, and push the result to a registry — each stage gating the next rather than running independently.

Why it is needed: skipping the order — for example, pushing before scanning, or tagging only after the fact — either lets an unscanned image reach a registry other systems might already be pulling from, or leaves you unable to answer "which exact image is running in production" when it matters most.

How it works, step by step (a standard GitHub Actions-based flow, using Docker's own official actions as a concrete, verifiable example):

  1. docker/setup-buildx-action configures the BuildKit-based builder, which is required for advanced features like SBOM and provenance generation that the default builder doesn't support.
  2. docker/metadata-action derives image tags and labels automatically from Git context — branch name, commit SHA, semantic version — so tags are consistent and traceable back to source without manual string-building in the workflow.
  3. docker/login-action authenticates to the target registry using credentials stored as CI secrets, never hardcoded in the workflow file.
  4. docker/build-push-action builds the image and, with a small amount of additional configuration, generates SBOM and provenance attestations at build time before pushing.
  5. A vulnerability scan (for example, with Docker Scout's CLI, checking for critical or high-severity CVEs) runs against the built image, with the pipeline configured to fail — stopping the push or a subsequent deployment step — if unacceptable vulnerabilities are found.

What fails without it: a pipeline missing the metadata-generation step tends to accumulate ad hoc, inconsistent tagging conventions across different team members' workflow edits. A pipeline missing the scan-and-gate step can push an image with known critical vulnerabilities straight to a registry other environments already pull from automatically.

✅ Worked example: Docker's own documentation shows a GitLab CI snippet that builds an image, then runs docker scout cves against it with --exit-code --only-severity critical,high, and pushes the image only if that scan step succeeds — a concrete, minimal illustration of scan-then-push as an enforced gate rather than an informational side step.

3. Build Caching — Why CI Builds Are Slow, and How to Fix It

🧸 Kid-friendly analogy: Re-chopping every vegetable from scratch each time you cook the same soup, instead of keeping pre-chopped ones in the fridge from last time, is exactly what an uncached CI build does to your Docker layers — redoing work that hasn't actually changed.

What it does: BuildKit can cache individual build layers and reuse them across builds, so steps unaffected by a given code change (installing dependencies that haven't changed, for instance) don't need to be redone.

Why it is needed: unlike a developer's own machine, where Docker's local build cache naturally persists between builds, CI runners are frequently ephemeral — a fresh environment on every run, with no local cache to reuse unless the pipeline explicitly exports and re-imports one.

How it works, step by step:

  1. Structure the Dockerfile so steps least likely to change (installing system packages, for example) come before steps most likely to change (copying application source code), so cache hits happen as early as possible in the build.
  2. Configure the CI build to export its cache to a location that persists between runs — a registry-backed cache, or a CI provider's own build-cache mechanism.
  3. On the next build, import that cache so unchanged layers are reused instead of rebuilt from scratch.
  4. Use multi-stage builds to keep the final image lean, separating build-time dependencies (compilers, build tools) from what's actually needed at runtime.

What fails without it: every CI run rebuilds every layer from scratch, which can turn what should be a two-minute incremental build into a fifteen-minute full rebuild — slow enough that teams start batching commits or skipping CI runs to avoid the wait, undermining the fast feedback CI is supposed to provide.

💡 Trade-off: registry-backed caching adds a small amount of pipeline complexity and registry storage cost, and cache correctness depends on your Dockerfile layer ordering actually reflecting what changes often versus rarely. A poorly ordered Dockerfile can make caching far less effective than expected, even with caching correctly configured at the pipeline level.

4. Worked Example — A GitHub Actions Pipeline from Commit to Registry (hypothetical, for illustration)

Consider a hypothetical team whose Docker pipeline originally just ran docker build and docker push, tagged only as latest.

How the pipeline evolved, step by step:

  1. They introduced docker/metadata-action to tag every image with both the Git commit SHA and a semantic version, ending the ambiguity of "which commit produced the image currently running in staging."
  2. They added a Docker Scout scan step configured to fail the pipeline on new critical or high-severity vulnerabilities, catching a vulnerable base-image update before it ever reached their registry.
  3. They enabled SBOM and provenance attestations in their build step, so a later security audit could answer "what's actually inside this image, and how was it built" without re-deriving that information by hand.
  4. They configured registry-backed build caching, cutting their typical CI build time from several minutes to under a minute for changes that didn't touch dependency files.

What this illustrates: none of these changes touched the application code itself — they were entirely pipeline-level improvements to how the image was built, verified, and versioned, and each one closed a specific gap (traceability, security, auditability, speed) rather than being adopted for its own sake.

🎯 Use this evolution as a rough adoption order: tagging and traceability first, then scanning, then attestation, then caching — front-load the changes that prevent the most costly mistakes.

5. Implementation — A Real Pipeline Definition

A representative GitHub Actions workflow combining tagging, SBOM/provenance attestation, and a build-and-push step:

# Illustrative example — verify current action versions against their own documentation
name: Build and Push Docker Image

on:
  push:
    branches: [main]
  pull_request:

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Set up Docker Buildx
        uses: docker/setup-buildx-action@v3

      - name: Extract metadata
        id: meta
        uses: docker/metadata-action@v5
        with:
          images: myorg/myapp

      - name: Log in to registry
        if: github.event_name != 'pull_request'
        uses: docker/login-action@v3
        with:
          username: ${{ secrets.REGISTRY_USERNAME }}
          password: ${{ secrets.REGISTRY_PASSWORD }}

      - name: Build and push
        uses: docker/build-push-action@v6
        with:
          context: .
          push: ${{ github.event_name != 'pull_request' }}
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          sbom: true
          provenance: true
          cache-from: type=gha
          cache-to: type=gha,mode=max

Illustrative example only — confirm current action versions, inputs, and cache backend options against each action's own documentation before use.

Two details in this pattern matter beyond the syntax:

  • Building on pull requests without pushing. Running the build (and ideally a scan) on every pull request, gated so it never pushes, catches build breakage and new vulnerabilities before merge — without ever putting an unreviewed image into a shared registry.
  • Enabling SBOM and provenance at build time, not as an afterthought. Generating these at build time captures accurate build parameters and source information directly, rather than trying to reconstruct that information from a finished image later.

6. Enterprise Rollout — Promotion, Rollback, and Supply-Chain Governance

Moving from "a pipeline that works" to "a pipeline the organization can rely on" adds governance concerns on top of the mechanics above:

  • Immutable, traceable tags. Tag every image with something that uniquely and permanently identifies its source — a commit SHA or content digest — rather than relying solely on a mutable tag like latest that can silently point to a different image over time.
  • Environment promotion, not environment rebuilding. Promote the exact same image (by digest) from staging to production rather than rebuilding from source at each stage — rebuilding risks introducing subtle differences between what was tested and what's deployed.
  • Defined rollback criteria. Because a previous image tag or digest is still available in the registry, rolling back is typically a matter of redeploying that reference — decide in advance what triggers this decision rather than improvising during an incident.
  • Vulnerability scan gates tied to policy, not vibes. Define explicitly which severities block a release (commonly critical and high) and where exceptions are tracked and time-boxed, rather than leaving the decision to whoever happens to be running the pipeline that day.
  • Supply-chain attestations as a compliance artifact. SBOMs and provenance records aren't just nice-to-have metadata — they're often the concrete evidence an audit or a customer's security review will ask for, and they're far cheaper to generate at build time than to reconstruct after the fact.
  • Registry access control. Restrict who can push to production-facing registry namespaces, and keep registry credentials as scoped CI secrets rather than long-lived personal credentials shared across a team.
  • Base image update cadence. A pipeline that pins base images by digest for reproducibility still needs a deliberate process for periodically updating those pins, or reproducibility becomes a way to freeze in place a base image with vulnerabilities that have since been fixed upstream.

7. Common Mistakes

  • Tagging only with latest. Without a commit-traceable tag, answering "which code produced the image currently running in production" requires guesswork, exactly when an incident makes that answer most urgent.
  • Pushing before scanning. An image pushed to a registry before its vulnerability scan completes can already be pulled by another system before the scan result comes back — the gate has to come before the push, not after.
  • Building without layer caching in CI. This turns every commit into a full rebuild, slowing feedback loops enough that teams start avoiding small, frequent commits — the opposite of what fast CI is meant to encourage.
  • Rebuilding the image at each deployment stage instead of promoting the same digest. Even with an identical Dockerfile and source commit, rebuilding can pull in a newer version of an unpinned dependency, meaning staging and production may not actually be running the same bits.
  • Treating SBOM and provenance generation as optional. Skipping this at build time means reconstructing "what's in this image" after the fact during a security incident or audit — far more expensive than generating it automatically as part of the build.
  • Storing registry credentials as long-lived, broadly-shared secrets. Credentials that don't expire and aren't scoped to a specific pipeline are a much larger blast radius if they ever leak than short-lived, narrowly-scoped CI secrets.

❓ FAQ

Should vulnerability scanning happen before or after pushing an image?

Before. If the scan runs after the push, other systems or environments could already have pulled the image by the time a critical vulnerability is discovered. The scan should gate the push, not follow it.

Why is tagging with just latest a problem?

Because latest is a mutable pointer that can silently refer to a different image over time, it makes it hard to know exactly which build is running anywhere, and complicates rollback since there's no fixed, previous reference to fall back to.

What are SBOM and provenance attestations, and why generate them at build time?

An SBOM lists exactly what's inside an image; provenance records how and from what source it was built. Generating both at build time captures accurate information directly, rather than trying to reconstruct it from a finished image later, which is harder and less reliable.

Should the same image be rebuilt for staging and production?

Generally no. Promoting the exact same image by tag or digest from staging to production avoids subtle differences that a rebuild could introduce, such as an unpinned dependency resolving to a newer version between builds.

How much does build caching actually help in CI?

It depends heavily on Dockerfile layer ordering, but for many projects it turns a full rebuild into a much faster incremental one whenever a change doesn't touch earlier, cached layers such as dependency installation steps.

🔗 References & Further Reading

"Docker," "Docker Scout," "Docker Hub," and related marks are trademarks of Docker, Inc.; "GitHub Actions" is a trademark of GitHub, Inc. 

📝 Summary

  • Docker image CI/CD deserves its own discipline because the image itself, not just the code inside it, is the deployable, auditable artifact.
  • A sound pipeline runs build, tag, scan, attest, and push in that order, each stage gating the next rather than running independently.
  • Build caching, done through registry-backed or CI-native cache mechanisms and sensible Dockerfile layer ordering, is what keeps CI feedback loops fast as a project grows.
  • A worked example shows a pipeline evolving from bare build-and-push into a traceable, scanned, attested workflow — one gap closed at a time.
  • A real pipeline definition shows how Docker's own official GitHub Actions handle tagging, authentication, and attestation with a small, composable set of steps.
  • Enterprise rollout adds immutable tagging, digest-based promotion, defined rollback criteria, and supply-chain attestations as genuine compliance artifacts.
  • Most real incidents trace back to a short list of avoidable mistakes: mutable-only tags, scanning after the push, uncached builds, and long-lived shared registry credentials.

Hopefully this gives you a clear, first-principles map of what a Docker image's journey from commit to registry should look like — and which gate to add first if your own pipeline is missing one. 🙌

Comments