Imagine you are baking a cake 🎂. Every time you change the recipe, you have to mix it yourself, put it in the oven yourself, and deliver it yourself. Every. Single. Time.
That sounds exhausting — especially if you change the recipe ten times a day!
CI/CD is the robot baker that does all of that for you automatically.
You change one line of code, push it to Git, and minutes later your new app is live on Kubernetes.
No manual steps. No human errors. No 3am deployments.
What is CI/CD?
CI/CD stands for Continuous Integration and Continuous Delivery.
These are two separate ideas that work together like a factory assembly line.
- CI (Continuous Integration) → Every time a developer pushes code, the pipeline automatically builds it, tests it, and packages it into a Docker image.
- CD (Continuous Delivery) → That Docker image is automatically deployed to Kubernetes — dev, staging, and production — with zero manual steps.
💡 Think of it like a post office:
You drop a letter in the box (push code).
The post office sorts, stamps, and delivers it automatically (CI/CD pipeline).
Your recipient gets it the same day (new version live on Kubernetes).
Together they mean: code written by a developer at 9am can be running in production by 9:05am — automatically and safely. ⚡
The Standard — GitOps
The industry has settled on one clear approach called GitOps.
The core idea is simple: Git is the single source of truth for everything.
Whatever is in your Git repository is exactly what runs in your cluster.
Want to deploy a new version? Don't run kubectl apply by hand — just push to Git.
Want to roll back? Just revert the Git commit. The cluster follows automatically.
- Application repo → Your Python/Go/Node.js code + Dockerfile
- GitOps repo → Your Kubernetes YAML files and Helm chart values
- CI tool → Builds the Docker image and updates the GitOps repo
- CD tool → Watches the GitOps repo and applies changes to the cluster
According to the CNCF Annual Survey 2025, 82% of container users run Kubernetes in production, and GitOps is now the dominant deployment strategy among them.
The Toolchain — What Everyone Is Using
The most popular CI/CD stack for Kubernetes looks like this:
- GitHub Actions → CI tool that builds, tests, and pushes your Docker image. Free for public repos.
- ArgoCD → CD tool that watches your GitOps repo and deploys to Kubernetes automatically.
- GitHub Container Registry (GHCR) → Stores your Docker images alongside your code. No separate registry needed.
- Trivy → Scans your Docker image for security vulnerabilities before it ever reaches production.
- Helm + Kustomize → Manages your Kubernetes manifests per environment.
Other valid choices in the ecosystem:
- GitLab CI → Best if your code lives on GitLab. Built-in CI/CD with zero extra setup.
- Jenkins → Still widely used in large enterprises with complex, multi-technology pipelines.
- FluxCD → An alternative to ArgoCD. Lighter weight, great for multi-cluster setups.
- Tekton → Cloud-native pipeline engine built directly on Kubernetes CRDs.
For this post we build the GitHub Actions + ArgoCD + GHCR stack — the most popular combination .
The Full Pipeline — How It All Fits Together
Here is the complete journey from a developer typing code to that code running in production:
Developer pushes code to GitHub
↓
GitHub Actions triggers automatically
↓
Step 1: Run unit tests
Step 2: Build Docker image
Step 3: Scan image with Trivy for vulnerabilities
Step 4: Push image to GHCR (tagged with commit SHA)
Step 5: Update image tag in GitOps repository
↓
ArgoCD detects the change in GitOps repo
↓
ArgoCD syncs the new manifest to Kubernetes
↓
Kubernetes performs a rolling update
Old pods replaced with new pods — zero downtime
↓
New version is live in production! ✅
The entire journey takes 3 to 8 minutes depending on your test suite and image size.
The developer just pushed code — they didn't touch a single server. 🎉
Part 1 — Setting Up GitHub Actions (CI)
GitHub Actions is a CI tool that runs automated workflows triggered by Git events.
You define your workflow in a YAML file inside .github/workflows/ in your repository.
Every push, pull request, or tag can trigger a different workflow.
Think of GitHub Actions like a row of dominoes 🁣.
You push code (tip the first domino) and every step falls automatically — test, build, scan, push, update.
Step 1: Repository Structure
Set up two separate Git repositories — one for app code, one for deployment config:
# Repository 1: weather-ai-app (application code)
weather-ai-app/
├── .github/
│ └── workflows/
│ └── ci.yaml ← your GitHub Actions pipeline
├── src/
│ └── app.py ← your Python ML model
├── tests/
│ └── test_model.py ← unit tests
├── Dockerfile ← how to build your container
└── requirements.txt
# Repository 2: weather-ai-gitops (deployment config)
weather-ai-gitops/
├── base/
│ ├── deployment.yaml ← K8s Deployment template
│ └── service.yaml
└── envs/
├── dev/
│ └── values.yaml ← dev-specific settings
└── production/
└── values.yaml ← production-specific settings
Keeping these repos separate is critical. It means a broken app build never corrupts your deployment config, and you get a clean audit trail of every deployment decision.
Step 2: Write the Dockerfile
Before CI can build your image, you need a Dockerfile that describes how to package your app:
# Dockerfile — multi-stage build
# Stage 1: Build dependencies
FROM python:3.12-slim AS builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir --user -r requirements.txt
# Stage 2: Final lean image — no build tools included
FROM python:3.12-slim
WORKDIR /app
COPY --from=builder /root/.local /root/.local
COPY src/ .
# Run as non-root user — security best practice
RUN adduser --disabled-password --gecos '' appuser
USER appuser
EXPOSE 8080
CMD ["python", "app.py"]
Why multi-stage?
- Stage 1 installs all the build tools (compilers, pip, etc.)
- Stage 2 copies only the final output — no build tools included
- Result: your production image is tiny, fast, and has a smaller attack surface
Step 3: Write the GitHub Actions CI Workflow
This is the heart of your CI pipeline. Create this file at .github/workflows/ci.yaml:
# .github/workflows/ci.yaml
name: CI — Build, Scan, and Push
on:
push:
branches: [ main ] # trigger on every push to main
pull_request:
branches: [ main ] # also run on pull requests
env:
REGISTRY: ghcr.io
IMAGE_NAME: ${{ github.repository }} # e.g. my-org/weather-ai-app
jobs:
test:
name: Run Unit Tests
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install dependencies
run: pip install -r requirements.txt
- name: Run tests
run: pytest tests/ --tb=short
build-and-push:
name: Build, Scan, and Push Image
runs-on: ubuntu-latest
needs: test # only run if tests passed!
permissions:
contents: write # needed to push to GitOps repo
packages: write # needed to push to GHCR
outputs:
image-tag: ${{ steps.meta.outputs.version }}
steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Generate image metadata and tags
id: meta
uses: docker/metadata-action@v5
with:
images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
tags: |
type=sha,prefix=sha- # e.g. sha-abc1234 — unique per commit
type=raw,value=latest,enable=${{ github.ref == 'refs/heads/main' }}
- name: Log in to GitHub Container Registry
uses: docker/login-action@v3
with:
registry: ${{ env.REGISTRY }}
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }} # auto-provided by GitHub
- name: Build Docker image
uses: docker/build-push-action@v5
with:
context: .
push: false # build but don't push yet — scan first!
tags: ${{ steps.meta.outputs.tags }}
load: true # load into local Docker daemon for scanning
- name: Scan image for vulnerabilities with Trivy
uses: aquasecurity/trivy-action@master
with:
image-ref: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}:${{ steps.meta.outputs.version }}
format: table
exit-code: 1 # FAIL the build if CRITICAL CVEs found
severity: CRITICAL,HIGH
- name: Push image to GHCR
uses: docker/build-push-action@v5
with:
context: .
push: true # now push — it passed the scan!
tags: ${{ steps.meta.outputs.tags }}
update-gitops:
name: Update GitOps Repository
runs-on: ubuntu-latest
needs: build-and-push # only run after successful push
if: github.ref == 'refs/heads/main' # only for main branch merges
steps:
- name: Checkout GitOps repository
uses: actions/checkout@v4
with:
repository: my-org/weather-ai-gitops
token: ${{ secrets.GITOPS_TOKEN }} # PAT with repo write access
- name: Update image tag in production values
run: |
# Replace the image tag in the Helm values file
sed -i "s|tag:.*|tag: sha-${{ github.sha }}|" \
envs/production/values.yaml
- name: Commit and push the updated manifest
run: |
git config user.name "github-actions[bot]"
git config user.email "github-actions[bot]@users.noreply.github.com"
git add envs/production/values.yaml
git commit -m "ci: update image to sha-${{ github.sha }}"
git push
That is your complete CI pipeline! Three jobs run in sequence: test → build/scan/push → update GitOps repo. ArgoCD takes over from here.
Part 2 — Setting Up ArgoCD (CD)
ArgoCD is a Kubernetes-native tool that watches your GitOps repo and makes your cluster match whatever is in Git.
Think of ArgoCD as a very strict security guard 👮.
If your cluster doesn't exactly match Git, ArgoCD fixes it automatically.
Step 4: Install ArgoCD on Your Cluster
# Create the argocd namespace
kubectl create namespace argocd
# Install ArgoCD using the official manifests
kubectl apply -n argocd \
--server-side \
--force-conflicts \
-f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml
# Wait for all ArgoCD pods to be running
kubectl wait --for=condition=Ready pod \
--all \
-n argocd \
--timeout=120s
# Get the initial admin password
kubectl get secret argocd-initial-admin-secret \
-n argocd \
-o jsonpath="{.data.password}" | base64 --decode
# Access the ArgoCD UI (forward port to your local machine)
kubectl port-forward svc/argocd-server -n argocd 8080:443
Open https://localhost:8080 in your browser. Log in with username admin and the password from above. You will see the ArgoCD dashboard — currently empty because we haven't told it what to deploy yet.
Step 5: Create an ArgoCD Application
An ArgoCD Application is a YAML object that tells ArgoCD: "watch this Git repo and keep this Kubernetes namespace in sync with it".
# argocd-app-production.yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: weather-ai-production
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/my-org/weather-ai-gitops.git
targetRevision: HEAD # always track the latest commit
path: envs/production # folder inside the GitOps repo
helm:
valueFiles:
- values.yaml # use production values
destination:
server: https://kubernetes.default.svc
namespace: production
syncPolicy:
automated:
prune: true # delete resources removed from Git
selfHeal: true # fix any manual changes made to the cluster
syncOptions:
- CreateNamespace=true
# Apply the ArgoCD Application to your cluster
kubectl apply -f argocd-app-production.yaml
# Watch ArgoCD sync in real time
kubectl get applications -n argocd -w
Expected output:
NAME SYNC STATUS HEALTH STATUS
weather-ai-production Synced Healthy
ArgoCD is now watching your GitOps repo. The moment GitHub Actions pushes a new image tag, ArgoCD sees the Git change and updates the cluster within seconds. ✅
What selfHeal means: If someone manually runs kubectl apply and changes something in the cluster, ArgoCD immediately reverts it back to match Git. Git is always the truth. Always.
Part 3 — The Complete End-to-End Flow
Let's trace one complete deployment from a code change to production, step by step.
Step 6: Make a Code Change and Push
# Developer makes a change to the model
echo "# Updated batch size to 64" >> src/app.py
# Stage and push the change
git add src/app.py
git commit -m "feat: increase default batch size to 64"
git push origin main
This single git push triggers everything automatically. Here is what happens next, all without any more human input:
00:00 — git push lands on GitHub
00:01 — GitHub Actions CI pipeline starts
00:15 — pytest runs 47 tests — all pass ✅
00:45 — Docker image built: sha-abc1234
01:20 — Trivy scans image — 0 critical CVEs found ✅
01:45 — Image pushed to ghcr.io/my-org/weather-ai-app:sha-abc1234
01:50 — GitOps repo updated: values.yaml tag → sha-abc1234
01:55 — ArgoCD detects the Git change
02:00 — ArgoCD syncs: new Deployment applied to production namespace
02:30 — Kubernetes rolling update: 5 new pods replacing 5 old pods
03:00 — All 5 new pods are Running and Ready ✅
03:00 — New version is live in production! 🎉
Total time: 3 minutes. Developer involvement after the push: zero. 🚀
Step 7: Verify the Deployment
# Check the ArgoCD app status
kubectl get applications -n argocd
# See the new pods running in production
kubectl get pods -n production
# Check the rollout history
kubectl rollout history deployment/weather-ai -n production
# See the exact image each pod is running
kubectl get pods -n production \
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[0].image}{"\n"}{end}'
Expected output:
NAME READY STATUS IMAGE
weather-ai-7d9f4-abc12 1/1 Running ghcr.io/my-org/weather-ai-app:sha-abc1234
weather-ai-7d9f4-def34 1/1 Running ghcr.io/my-org/weather-ai-app:sha-abc1234
weather-ai-7d9f4-ghi56 1/1 Running ghcr.io/my-org/weather-ai-app:sha-abc1234
All pods running the brand new image. Deployed automatically. Clean! 🎯
Part 4 — Rolling Back When Things Go Wrong
Deployments can fail. A new model might return wrong predictions. A bug might slip through testing.
With GitOps, rolling back is as simple as reverting a Git commit.
Option 1: Git Revert (GitOps Way — Recommended)
# Go to the GitOps repository
cd weather-ai-gitops
# Find the previous good commit
git log --oneline envs/production/values.yaml
# Revert to the previous commit
git revert HEAD
# Push the revert — ArgoCD detects and redeploys the old version
git push origin main
ArgoCD sees the reverted tag, syncs immediately, and your old version is running within 60 seconds. The rollback itself is a new Git commit — fully auditable. 📝
Option 2: ArgoCD CLI Rollback (Emergency Fast Path)
# Install ArgoCD CLI
curl -sSL -o /usr/local/bin/argocd \
https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
chmod +x /usr/local/bin/argocd
# Log in
argocd login localhost:8080 --username admin --insecure
# See the deployment history for your app
argocd app history weather-ai-production
# Roll back to a specific revision
argocd app rollback weather-ai-production 3
# Verify the status after rollback
argocd app get weather-ai-production
⚠️ Important: ArgoCD CLI rollback is temporary. Because selfHeal: true is set, ArgoCD will eventually re-sync from Git and overwrite the rollback. Always follow up with a proper Git revert to make the rollback permanent.
Part 5 — Security Scanning with Trivy
IEvery CI pipeline must include a vulnerability scan before any image reaches production.
Trivy is the most widely adopted open-source scanner — it checks your image for known security vulnerabilities in base OS packages, Python libraries, Node modules, and more.
We already included Trivy in our GitHub Actions workflow above. Here is how to run it locally too:
# Install Trivy on Mac
brew install trivy
# Install Trivy on Linux
curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh
# Scan your image — show only HIGH and CRITICAL issues
trivy image \
--severity HIGH,CRITICAL \
ghcr.io/my-org/weather-ai-app:sha-abc1234
Expected output (clean image):
ghcr.io/my-org/weather-ai-app:sha-abc1234 (debian 12.4)
============================================================
Total: 0 (HIGH: 0, CRITICAL: 0)
If Trivy finds issues:
Total: 3 (HIGH: 2, CRITICAL: 1)
Library Vulnerability Severity Fixed Version
─────────────────────────────────────────────────────────
libssl3 CVE-2024-12345 CRITICAL 3.0.13 ← update base image!
cryptography CVE-2024-67890 HIGH 42.0.5 ← upgrade in requirements.txt
requests CVE-2024-11111 HIGH 2.32.0 ← upgrade in requirements.txt
The CI pipeline fails here — the image is never pushed to GHCR. Fix the vulnerabilities, push again, and the pipeline retries.
⚠️ Trend: Companies are now scanning for Software Bill of Materials (SBOM) too. An SBOM is a complete list of every library inside your container — like a nutritional label for software. Generate one with:
# Generate SBOM for your image
trivy image \
--format spdx-json \
--output sbom.json \
ghcr.io/my-org/weather-ai-app:sha-abc1234
Part 6 — Progressive Delivery with Argo Rollouts
A standard Kubernetes rolling update replaces all pods with the new version all at once.
If the new version has a bug, all your users are affected instantly.
Progressive delivery releases the new version gradually — to a small percentage of users first, checks metrics, then widens the rollout if everything looks good. Think of it like launching a rocket 🚀 — you test the engines at 10% thrust before going to 100%.
Argo Rollouts is the standard tool for this on Kubernetes.
Canary Deployment — New Version to 10% of Users First
# argo-rollout.yaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: weather-ai
namespace: production
spec:
replicas: 10
strategy:
canary:
steps:
- setWeight: 10 # send 10% of traffic to new version
- pause: {duration: 5m} # wait 5 minutes and monitor
- setWeight: 30 # increase to 30%
- pause: {duration: 5m}
- setWeight: 60 # increase to 60%
- pause: {duration: 5m}
- setWeight: 100 # full rollout — all users on new version
# Automatic rollback if error rate goes above 5%
analysis:
templates:
- templateName: error-rate-check
args:
- name: service-name
value: weather-ai
selector:
matchLabels:
app: weather-ai
template:
spec:
containers:
- name: weather-ai
image: ghcr.io/my-org/weather-ai-app:sha-abc1234
With this setup, if error rates spike during the canary phase, Argo Rollouts automatically aborts the deployment and rolls back to the stable version — no human needed. ✅
Blue-Green Deployment — Zero-Downtime Instant Switch
# Blue-Green strategy (alternative to canary)
strategy:
blueGreen:
activeService: weather-ai-active # current live version (blue)
previewService: weather-ai-preview # new version (green) — not live yet
autoPromotionEnabled: false # require manual promotion
scaleDownDelaySeconds: 30 # keep old version warm for 30 seconds after switch
- Canary → Gradual traffic shift. Best for catching performance issues at scale.
- Blue-Green → Two full environments. Instant switch. Best for zero-downtime releases with quick rollback.
Part 7 — Multi-Environment Pipeline (Dev → Staging → Production)
Real MLOps teams always deploy through multiple environments before production.
Here is the full GitOps folder structure for managing three environments:
weather-ai-gitops/
├── base/
│ ├── deployment.yaml ← shared base template
│ └── service.yaml
└── envs/
├── dev/
│ └── values.yaml ← dev: replicas=1, debug logging, latest tag
├── staging/
│ └── values.yaml ← staging: replicas=2, info logging, rc tag
└── production/
└── values.yaml ← production: replicas=10, warn logging, sha tag
Each environment has its own ArgoCD Application pointing to its folder:
# Create ArgoCD apps for all three environments at once
for env in dev staging production; do
kubectl apply -f - <<EOF
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: weather-ai-${env}
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/my-org/weather-ai-gitops.git
targetRevision: HEAD
path: envs/${env}
helm:
valueFiles: [values.yaml]
destination:
server: https://kubernetes.default.svc
namespace: ${env}
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
EOF
done
Now each environment auto-syncs independently. Push to dev. Test. Promote to staging. Test again. Promote to production. All via Git commits — no kubectl, no manual server access, no risk of human error.
Essential Commands Cheat Sheet
GitHub Actions — Useful Commands
# Trigger a workflow manually from CLI (using GitHub CLI)
gh workflow run ci.yaml --ref main
# View workflow run logs
gh run list --workflow=ci.yaml
gh run view <run-id> --log
# View recent workflow runs in browser
gh run list
ArgoCD — Useful Commands
# List all ArgoCD applications
argocd app list
# Get detailed status of one app
argocd app get weather-ai-production
# Manually trigger a sync (force ArgoCD to re-check Git now)
argocd app sync weather-ai-production
# View deployment history
argocd app history weather-ai-production
# Roll back to a specific revision
argocd app rollback weather-ai-production <revision-number>
# Diff: what would change if ArgoCD synced right now?
argocd app diff weather-ai-production
Trivy Security Scanning
# Scan image for vulnerabilities
trivy image --severity HIGH,CRITICAL <image>:<tag>
# Scan a filesystem (scan before building the image)
trivy fs --severity HIGH,CRITICAL .
# Generate SBOM
trivy image --format spdx-json --output sbom.json <image>:<tag>
# Scan Kubernetes manifests for misconfigurations
trivy k8s --report summary cluster
Argo Rollouts
# Install Argo Rollouts
kubectl create namespace argo-rollouts
kubectl apply -n argo-rollouts \
-f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml
# Watch a canary rollout progress in real time
kubectl argo rollouts get rollout weather-ai -n production --watch
# Manually promote a paused canary to the next step
kubectl argo rollouts promote weather-ai -n production
# Abort a canary and roll back immediately
kubectl argo rollouts abort weather-ai -n production
Quick Summary 📝
What we learned today:
- What CI/CD is → Automated pipeline from code push to production deployment. No manual steps.
- GitOps → Git is the single source of truth. Git changes drive cluster changes. Auditable, reversible, safe.
- GitHub Actions → CI tool. Runs tests, builds Docker image, scans for vulnerabilities, pushes to GHCR, updates GitOps repo.
- ArgoCD → CD tool. Watches GitOps repo. Auto-syncs Kubernetes cluster to match Git. Self-heals drift.
- Trivy → Security scanner. Fails the build before CRITICAL vulnerabilities ever reach production.
- Argo Rollouts → Progressive delivery. Canary and Blue-Green strategies. Automatic rollback on bad metrics.
- Multi-environment → Dev → Staging → Production all managed by separate ArgoCD apps pointing to separate folders.
- Rollback → Git revert triggers automatic redeployment of the old version. Full audit trail in Git history.
Keep building & automating! ⚙️✨
Comments
Post a Comment