Skip to main content

Posts

Showing posts with the label Docker

AI Model Serving with Kubernetes: Deploy, Scale, and Manage ML Models

AI model serving on Kubernetes means running inference workloads under a scheduler and autoscaler that were originally designed for stateless web services, then deliberately adapting both — GPU-aware scheduling, request-based rather than CPU-based autoscaling, and slower, heavier "pod startup" — to fit a workload where loading a model into memory can take longer than most web requests take to complete. 🧠 A standard Kubernetes Deployment with a CPU-based Horizontal Pod Autoscaler will technically run a model server. It just won't run it well: it can't tell the difference between an idle GPU and a busy one, it scales on the wrong signal, and it has no concept of a canary rollout for a new model version versus a new code version. Purpose-built model-serving platforms on Kubernetes — KServe chief among them — exist specifically to close that gap. Skipping them doesn't mean model serving becomes simpler; it means you end up quietly rebuilding pieces of them yours...

Production Releases and Rollbacks: Safe Docker Deployment Strategies

A production release strategy is the controlled process for replacing running containers with a new version without breaking the service in front of them, and rollback is the equally controlled process for undoing that change the moment something goes wrong. In Kubernetes, both are built on the same underlying mechanism — the Deployment controller gradually reconciling actual state toward a declared desired state — which means understanding that one mechanism explains how releases, rollbacks, and canaries all actually work under the hood.  The stakes here are immediate and visible to users in a way few other topics in this series are. A botched deployment doesn't fail quietly in a log file somewhere — it shows up as errors in someone's browser, right now, while a team scrambles to figure out whether to keep pushing forward or pull the rip cord. Knowing exactly what a rollback command actually does, and how fast it can act, is the difference between a five-minute blip and a...

CI/CD for Docker Images: Automate Build, Testing, and Deployment

CI/CD for Docker images is the automated pipeline that turns a Dockerfile change into a scanned, tagged, attested, and pushed image ready for deployment — and doing it well means treating the image itself, not just the application code inside it, as the artifact that needs testing, versioning, and security gates. 🏗️ Teams that bolt Docker onto an existing CI/CD pipeline as an afterthought tend to hit the same wall: builds that take twenty minutes because nothing is cached, a latest tag that makes rollback guesswork, and a vulnerability discovered in production that a five-second scan would have caught before it ever shipped. None of this is exotic — it's a small, well-understood set of pipeline stages, done in the right order, with the right gates. Get the order wrong, and you either slow every developer down or ship risk straight into production. 🚦 📑 In This Post Foundations: the inner loop, the outer loop, and why images need their own pipeline discipline Pipeline ...