AI model serving on Kubernetes means running inference workloads under a scheduler and autoscaler that were originally designed for stateless web services, then deliberately adapting both — GPU-aware scheduling, request-based rather than CPU-based autoscaling, and slower, heavier "pod startup" — to fit a workload where loading a model into memory can take longer than most web requests take to complete. 🧠 A standard Kubernetes Deployment with a CPU-based Horizontal Pod Autoscaler will technically run a model server. It just won't run it well: it can't tell the difference between an idle GPU and a busy one, it scales on the wrong signal, and it has no concept of a canary rollout for a new model version versus a new code version. Purpose-built model-serving platforms on Kubernetes — KServe chief among them — exist specifically to close that gap. Skipping them doesn't mean model serving becomes simpler; it means you end up quietly rebuilding pieces of them yours...