Model and dataset storage is the set of decisions and mechanisms that determine where large, mutable ML artifacts — trained weights, tokenizer files, embeddings indices, and training/evaluation datasets — physically live outside a container image, and how running Docker containers or Kubernetes pods reach them reliably, securely, and fast. Unlike application code, model weights and datasets are large (often gigabytes to hundreds of gigabytes), change independently of the code that serves them, and are frequently shared across many replicas at once. 📦
Getting this wrong is not cosmetic. Baking a 14 GB model into an image bloats every build and every node's disk; a missing shared volume turns autoscaling into a stampede of duplicate downloads; and an unversioned dataset silently invalidates weeks of evaluation history. In production AI infrastructure, storage architecture is as consequential as the model architecture itself — it decides your rollout speed, your blast radius during a bad deploy, and whether a "which dataset produced this model" question has an answer six months from now. ⚙️
Figure 1 — Original diagram: the image stays thin; weights and datasets travel through storage chosen for how widely they need to be shared.
📑 In This Post
- 1. Foundations: What "Model and Dataset Storage" Actually Covers
- 2. Mechanics: Image Layers, Volumes, Bind Mounts, and Caches
- 3. Worked Example: A Hugging Face Model from Docker to Kubernetes
- 4. Implementation: Docker Compose and Kubernetes PVC Patterns
- 5. Versioning and Registries: DVC, MLflow, and OCI Model Artifacts
- 6. Enterprise Rollout: Governance, Access, Cost, and Incident Response
- 7. Common Mistakes
- 8. FAQ
- 9. References & Further Reading
- 10. Summary
🔀 Quick Comparison: Where Should the Bytes Live?
| Storage Pattern | Persists Past Container Restart? | Shared Across Pods/Nodes? | Update Without Rebuild? | Typical Fit |
|---|---|---|---|---|
| Baked into image | Yes (part of image) | Yes, but duplicated per pull | No — needs rebuild | Tiny models, edge/air-gapped deploys |
| Bind mount / local cache | Yes, on that host | No — single host | Yes | Local dev, single-node inference |
| PVC — ReadWriteOnce | Yes | No — one node at a time | Yes | Single-replica serving, training jobs |
| PVC — ReadWriteMany / object storage | Yes | Yes — many pods concurrently | Yes | Autoscaled multi-replica inference |
| Model registry / OCI artifact | Yes, versioned | Yes — pulled per node/pod | Yes, new version = new pull | Governed promotion, audit trail, CI/CD gates |
🎯 Use this when choosing how a specific workload should reach its weights before you write a single YAML manifest.
1. Foundations: What "Model and Dataset Storage" Actually Covers
Kid-friendly analogy: Think of a food truck. The truck itself (small, drives anywhere) is your Docker image — it carries the stove, the recipe cards, the cash register. But the truck does not carry a warehouse of ingredients inside it; it drives to a supply depot each morning and loads what it needs. The depot is your model and dataset storage: bigger, slower-moving, and shared by many trucks.
Translated technically: a Docker image is a stack of read-only layers assembled at build time and identified by a content hash. Model weights and datasets are usually orders of magnitude larger than application code, are produced by a separate process (training, fine-tuning, data collection) on a separate cadence, and frequently need to be read by many containers at once. Treating them as "just another file to COPY into the image" collapses two lifecycles — code release and data/model release — into one, which is the root cause of most storage-related production problems in ML infrastructure.
There are three broad ownership questions every model/dataset storage decision answers:
- Where does the artifact physically persist — inside the image, on a single host's disk, on a Kubernetes volume, or in an external object store or registry?
- Who else needs concurrent access — one process, one node's worth of pods, or every replica across the whole cluster?
- How is a specific version retrieved deterministically — by file path, by content hash, by a registry version number, or by a mutable "latest" pointer that can silently change under you?
Everything else in this post — bind mounts, PersistentVolumeClaims, Hugging Face caching, DVC, MLflow's Model Registry, OCI model artifacts — is a different answer to those same three questions, optimized for a different point in the ML lifecycle (local development, single-node inference, distributed training, autoscaled serving, or governed production promotion).
2. Mechanics: Image Layers, Volumes, Bind Mounts, and Caches
Kid-friendly analogy: A Docker image layer is like a sealed page in a book — once printed, it cannot be edited, only replaced by printing a new page. A volume is more like a sticky note attached to the book: you can scribble on it, swap it out, or share it with a friend's copy of the same book, without ever reprinting a page.
What it does. Docker builds an image as a sequence of immutable, content-addressed layers. A container adds one thin writable layer on top at runtime. Anything written there disappears when the container is removed unless it is explicitly persisted through a bind mount (a host directory mapped into the container) or a named volume (storage Docker manages on the host or via a plugin, independent of any single container's lifecycle). Kubernetes extends the same idea across a cluster with PersistentVolumes (PVs, the actual storage resource) and PersistentVolumeClaims (PVCs, a namespaced request for that storage), each carrying an accessMode.
Why it is needed. Model weights change on a training cadence, not a code-release cadence; datasets are frequently too large to fit inside a reasonable image and too sensitive to bake into something pushed to a registry with broad pull access. Separating "where the code lives" from "where the data lives" lets each scale, version, and be secured independently.
How it works, step by step:
- The image build produces layers for the OS, the CUDA/Python runtime, and the application code — no weights.
- At container or pod start, the orchestrator resolves the requested volume: a bind mount path, a Docker named volume, or a Kubernetes PVC bound to a PV.
- An init step (a first-run cache check, or in Kubernetes an explicit init container) checks whether the artifact is already present at the mount path.
- If absent, it is pulled from its source of truth — an object store, a Hugging Face repository, or a model registry — and written into the mounted path, not into the container's writable layer.
- The main process opens the model or dataset files directly from the mount, the same way it would open any local file.
- On restart, if the volume persisted, step 3 short-circuits and startup is fast; if it did not, the full download repeats.
What fails without it. Without a persistent mount, every container restart re-downloads the full model — multiplying egress cost and startup latency, and turning a rolling deployment or a crash-loop into a self-inflicted denial-of-service against the model source. Without the right access mode, two pods scheduled on different nodes may fail to mount the same ReadWriteOnce volume at all, producing pods stuck in ContainerCreating.
Hugging Face's own caching layer is a concrete illustration of this mechanic. The huggingface_hub library resolves a cache root from, in order, an explicit cache_dir argument, the HF_HUB_CACHE environment variable, or $HF_HOME/hub, defaulting to ~/.cache/huggingface/hub. According to Hugging Face's own environment-variable reference, this setting specifically governs where downloaded models, datasets, and Spaces repositories are cached on disk, distinct from where tokens and logs live. Pointing that variable at a mounted volume, instead of leaving it at its container-local default, is the difference between every pod redownloading a multi-gigabyte model and every pod reusing one already-warmed cache.
✅ Practical example. A FastAPI inference service's Dockerfile sets ENV HF_HOME=/models/hf-cache, and the Kubernetes Deployment mounts a PVC at /models. The first pod to start pulls the model once; every subsequent restart of that pod — and, if the PVC's access mode allows it, every other pod sharing the volume — reads straight from disk with zero network calls to the model source.
💡 Trade-off. A shared cache is a shared failure domain too: a corrupted or partially-written cache entry (for example, from a pod killed mid-download) can be silently reused by every pod that mounts the same volume afterward. Production caches need integrity checks — content-hash verification, not just "file exists" — before trusting a cached artifact.
3. Worked Example: A Hugging Face Model from Docker to Kubernetes
Consider a team serving an open-weight model pulled from the Hugging Face Hub through a FastAPI wrapper. This example follows the documented mechanics of Hugging Face's own caching system and KServe's storage handling, which are publicly specified even though the exact team and traffic numbers below are a labeled hypothetical.
What it does. On first request, the service loads the model into memory; the underlying huggingface_hub client checks its cache directory before making any network call, and only downloads files that are missing or whose revision has changed.
Why it is needed. A cold pull of a multi-gigabyte model on every pod start directly taxes deploy speed, autoscaling responsiveness, and the model host's bandwidth. Reusing a warmed cache turns a multi-minute cold start into a near-instant one.
How it works, step by step (local Docker development):
- The Dockerfile installs transformers and huggingface_hub but does not download any weights during the build.
- docker run -v hf-cache:/root/.cache/huggingface ... attaches a named volume at the library's default cache path.
- On first start, the app calls from_pretrained(), which downloads model files into that volume, laid out as content-addressed blobs with a symlinked snapshot directory per revision.
- On the next docker run with the same volume, the same call finds the cached blobs and skips the download entirely.
How it works, step by step (Kubernetes production):
- An init container, or KServe's storage-initializer pattern, resolves a hf://-style URI and downloads the model into a mounted volume before the serving container starts, using KServe's documented ClusterStorageContainer mechanism for Hugging Face Hub sources.
- A shared token, stored as a Kubernetes Secret and injected as HF_TOKEN, authenticates gated model access, following the same secret-reference pattern documented for KServe's Hugging Face storage container.
- The serving container mounts the same volume read-only, so it never needs the download credentials itself — a useful blast-radius reduction.
- When scaling to multiple replicas, the access mode of the underlying volume decides whether every replica shares one cached copy or each triggers its own download; KServe's local-model-cache feature explicitly targets this by pre-staging models on node-local storage per node group.
✅ Practical example. A labeled hypothetical: a team running 12 GPU-backed replicas of a 13B-parameter model switches from "each pod downloads on start" to a pre-warmed node-local cache. Rolling restarts that previously took several minutes per pod, gated by download speed, instead complete in the time it takes the process to load already-local weights into GPU memory.
🎯 Use this when you are moving a Hugging Face–based service from a laptop to a multi-replica Kubernetes deployment and need predictable, fast cold starts.
4. Implementation: Docker Compose and Kubernetes PVC Patterns
Kid-friendly analogy: Choosing an access mode is like choosing a locker style: a gym locker one person opens (ReadWriteOnce), a library shelf many people can read but not rewrite (ReadOnlyMany), and a shared whiteboard many people can both read and write on at once (ReadWriteMany). Picking the wrong locker for the job either blocks people who needed in, or lets too many hands scribble at once.
Local development with Docker Compose. The goal here is simply to avoid redownloading on every rebuild:
# docker-compose.yml (illustrative example)
services:
inference:
build: .
environment:
- HF_HOME=/models/hf-cache
volumes:
- model-cache:/models/hf-cache # named volume, survives rebuilds
ports:
- "8000:8000"
volumes:
model-cache:
Rebuilding the image (code change) leaves model-cache untouched; deleting the volume forces a fresh pull, which is exactly the control you want when the underlying model revision changes.
Kubernetes: single-replica serving with a ReadWriteOnce PVC. This is the simplest production pattern and fits most low-QPS or early-stage services:
# model-pvc.yaml (illustrative example)
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: model-store
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 50Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: inference-svc
spec:
replicas: 1
template:
spec:
initContainers:
- name: fetch-model
image: myorg/model-fetcher:1.0
env:
- name: HF_HOME
value: /models/hf-cache
volumeMounts:
- name: model-vol
mountPath: /models
containers:
- name: server
image: myorg/inference:1.0
env:
- name: HF_HOME
value: /models/hf-cache
volumeMounts:
- name: model-vol
mountPath: /models
volumes:
- name: model-vol
persistentVolumeClaim:
claimName: model-store
The init container pattern is deliberate: it separates "fetch credentials and network access to the model source" from "serve requests," so the serving container's attack surface does not need outbound access to the model registry at all.
Kubernetes: multi-replica autoscaled serving. Because ReadWriteOnce binds a volume to one node, scaling replicas across nodes with that access mode either serializes scheduling onto a single node or leaves new pods unable to mount the claim at all. Kubernetes documentation for the ReadWriteMany mode is explicit that it exists precisely so many nodes can access the same volume simultaneously. In practice this means either provisioning a ReadWriteMany-capable storage class (an NFS-backed class or a cloud filesystem service), or moving to node-local caching with a controller that pre-stages the model on every node — the pattern KServe implements with its LocalModelNodeGroup resource, which pins a storage limit and node affinity for a local NVMe-backed cache shared by pods scheduled to that node.
💡 Trade-off. ReadWriteMany storage classes (NFS, cloud filesystem services) trade raw throughput for shareability; large models loaded from network filesystems can be slower to mmap into GPU memory than models on node-local NVMe. Benchmarking model-load latency under your actual storage class, not just download time, is part of a real capacity plan.
🎯 Use this when deciding between a simple single-replica PVC and a shared, multi-node caching layer as traffic and replica count grow.
5. Versioning and Registries: DVC, MLflow, and OCI Model Artifacts
Kid-friendly analogy: A folder full of files named "model_final.pt", "model_final_v2.pt", and "model_final_v2_REALLY.pt" is like a diary with torn-out, undated pages. A model registry is a librarian who stamps every page with a date, a shelf number, and a note on whether it is a draft or the one currently checked out to production.
What it does. A storage location tells you where bytes live; a versioning system tells you which bytes they are and how they relate to the code and data that produced them. Three complementary approaches show up repeatedly in containerized ML stacks:
- DVC (Data Version Control) tracks large datasets and model files alongside Git by storing a small metafile — a content hash and path — in the Git repo while the actual bytes live in a configured remote (S3, GCS, Azure, SSH, or a local path). This keeps Git fast while still giving every dataset version a reproducible pointer tied to a specific commit.
- MLflow's Model Registry provides a centralized model store with APIs and a UI for versioning models, assigning them monotonically increasing version numbers, and attaching mutable aliases or lifecycle stages so a serving system can resolve a name like "Production" to whichever concrete version currently holds that label.
- OCI model artifacts, the approach Docker's Model Runner uses, package a model file (commonly in GGUF format) as an OCI artifact and push it to any OCI-compliant registry — reusing the same tagging, pulling, and access-control machinery teams already run for container images, rather than inventing a parallel distribution system.
Why it is needed. Without a registry or a hashed pointer, "which dataset trained this model" and "which model is actually running in production" become tribal knowledge, and rollback means guessing which file on someone's laptop was the last known-good one.
How it works, step by step (DVC dataset flow):
- dvc add dataset.parquet moves the file into DVC's local cache and writes a small .dvc metafile containing its hash.
- The metafile — not the dataset — is committed to Git, so history stays lightweight.
- dvc push uploads the actual bytes to a configured remote.
- Anyone who checks out that Git commit and runs dvc pull retrieves the exact dataset version tied to that commit — including inside a CI job or a training container.
# Illustrative DVC workflow dvc remote add -d storage s3://my-bucket/dvcstore dvc add data/train.parquet git add data/train.parquet.dvc .gitignore git commit -m "Track training set v3" dvc push
What fails without versioning. An evaluation run compares two model checkpoints, but the "old" checkpoint was quietly overwritten in the shared object store bucket the week before — the comparison is now meaningless, and nobody can tell because the file path never changed. Content-hashed, registry-tracked artifacts make that class of silent corruption structurally harder, because a changed file produces a new hash or a new version number instead of overwriting the old one in place.
✅ Practical example. A CI pipeline that promotes a model registers it in MLflow's registry with a new version number, then moves an alias such as "champion" to point at that version only after evaluation gates pass — so the serving layer's model reference (models:/fraud-detector@champion) never has to change, only what it resolves to.
🎯 Use this when you need a defensible answer to "exactly which weights and which data are running right now, and what shipped before them."
6. Enterprise Rollout: Governance, Access, Cost, and Incident Response
Moving model and dataset storage from "it works on the team's cluster" to an organization-wide standard means treating it as infrastructure with owners, not a convention. The areas below are the ones that consistently separate a storage pattern that scales from one that becomes an incident source.
Ownership and dataset/model versioning. Assign a clear owner for the registry or storage remote itself (who can create new namespaces, who approves new remotes) separately from who owns individual model or dataset entries. Require that every training or fine-tuning job records the exact dataset version (a DVC commit hash, a dataset registry version) and code commit it used, so a model version is always traceable back to both inputs.
CI gates. Treat a model registry promotion — moving an alias like "staging" to "production" — as a deploy, not a metadata edit. Gate it behind the same kind of checks used for code: automated evaluation thresholds, a required approval, and a recorded diff against the previous production version's metrics.
Access controls and data privacy. Datasets derived from production traffic often carry the same sensitivity as the production system itself — sometimes more, once they are aggregated and exported. Scope registry and bucket permissions by dataset classification, not by team boundary alone, and keep production-derived data out of default-readable buckets that every service account can reach.
Budget controls. Large shared caches and ReadWriteMany volumes are easy to over-provision and easy to forget. Track storage cost per model family and per dataset lineage the same way compute cost is tracked, and set retention policies so superseded checkpoints and DVC cache entries expire instead of accumulating indefinitely.
Dashboards and alerts. At minimum, alert on: cache-miss rate spikes (a sign the shared cache stopped being shared, or a new model revision is thrashing it), PVC utilization approaching capacity, and download failures from the model source distinguished from download failures due to storage-side quota or permission issues — these have very different remediation paths.
Incident response. Write a runbook for "the shared model cache is corrupted or stale" separately from "the model registry is unreachable." The first usually means invalidating and re-warming a specific cache path; the second usually means falling back to a known-good pinned artifact already resident on nodes, which is only possible if node-local pre-staging (as in KServe's local model cache pattern) was already in place before the incident.
💡 Trade-off. Node-local pre-staging gives the fastest, most resilient cold starts but duplicates storage per node and adds an extra controller and job namespace to operate and secure — it is worth the operational overhead only once replica count and blast-radius risk justify it.
🎯 Use this checklist when a storage pattern that worked for one team's cluster is being adopted as an org-wide standard.
7. Common Mistakes
Baking large models into the image "for simplicity." This seems fine at first because it removes a moving part, but every image pull now carries the full model weight, every build takes longer, node disk pressure increases with every model version ever pushed, and a one-line code fix forces a multi-gigabyte re-push. The causal chain runs from "image size" straight to "deploy latency" and "node storage exhaustion" under autoscaling.
Choosing ReadWriteOnce for a multi-replica, multi-node service. It works flawlessly in testing with one replica on one node, then fails unpredictably the moment the scheduler places a second replica on a different node, because the access mode was never designed to be shared across nodes — the result is pods stuck pending, not a graceful error.
Treating a shared cache as inherently safe to trust. A pod killed mid-download can leave a partially written file at the expected path; a naive existence check ("file is there, skip download") then serves a truncated or corrupted model to every subsequent pod that mounts the same cache. The production impact is a model that loads without error but produces garbage or crashes deep inside inference, which is far harder to diagnose than a failed download.
Overwriting dataset files in place instead of versioning them. Reusing the same object-store path for every new dataset export destroys the ability to reproduce or audit any earlier training run, and any evaluation comparison against "the old dataset" silently becomes invalid the moment that path is overwritten.
Giving the serving container the same credentials as the download step. Bundling model-source credentials into the long-running serving process, instead of isolating them in a short-lived init container, needlessly widens the blast radius if that container is ever compromised — the serving process never needed outbound access to the model registry in the first place.
No retention policy on caches or checkpoints. Every superseded model revision and every DVC cache entry that is never garbage-collected quietly grows the storage bill and, eventually, fills the volume — at which point new deployments start failing for a reason that has nothing to do with the code being shipped that day.
❓ FAQ
Should I ever put a model file inside a Docker image?
Sometimes — small models, edge deployments with no reliable network to a model source, or air-gapped environments genuinely benefit from a self-contained image. The trade-off is that every version bump means rebuilding and re-pushing the full image, so it fits best when the model changes rarely.
What is the practical difference between a Docker named volume and a Kubernetes PVC?
A Docker named volume is scoped to a single Docker host; a Kubernetes PersistentVolumeClaim is a cluster-level request that Kubernetes binds to an actual PersistentVolume, which can be backed by node-local disk, a networked filesystem, or a cloud storage service, and which carries an access mode that governs whether multiple nodes can use it at once.
Do I need DVC and a model registry, or just one of them?
They solve related but distinct problems: DVC is commonly used to version the datasets and files that feed a training pipeline alongside Git history, while a model registry like MLflow's tracks the trained model artifacts themselves, their versions, and their promotion status. Many production stacks use both, connected by recording the dataset version used for each registered model.
Why would I use ReadWriteMany instead of just giving every pod its own ReadWriteOnce volume?
Per-pod volumes mean per-pod downloads, which multiplies egress traffic and cold-start time with every added replica. A shared ReadWriteMany volume, or a node-local pre-staged cache, lets many pods read one already-warmed copy instead, at the cost of needing storage infrastructure that supports concurrent multi-node access.
Is this article eligible for a rich FAQ result in search engines?
This post includes FAQPage structured data that matches the questions and answers shown above. Whether any given search engine chooses to display a rich result is decided entirely by that search engine, varies over time, and is not something this article — or any publisher — can guarantee.
🔗 References & Further Reading
- Hugging Face Hub — Environment Variables reference (HF_HOME, HF_HUB_CACHE)
- Hugging Face Datasets — Cache Management documentation
- KServe — Local Model Cache documentation
- Kubernetes Blog — Introducing Single Pod Access Mode for PersistentVolumes
- NVIDIA NeMo Microservices — ReadWriteMany Persistent Volumes documentation
- NVIDIA Triton Inference Server — Conceptual Guide, Model Repository
- MLflow — Model Registry documentation
- DVC — Get Started guide (data and model versioning)
- Docker — Model Runner documentation (OCI model artifacts)
Product and project names above (Docker, Kubernetes, Hugging Face, KServe, NVIDIA Triton, MLflow, DVC) are trademarks of their respective owners and are used here only to identify the referenced technologies.
📝 Summary
- Foundations: model/dataset storage separates the code lifecycle from the data lifecycle, and every choice answers where bytes persist, who shares them, and how a version is pinned.
- Mechanics: immutable image layers plus mutable mounted storage — bind mounts, named volumes, or Kubernetes PVCs — keep images thin and let weights and data update independently of code.
- Worked example: a Hugging Face–backed service moves from a local Docker cache to an init-container-driven Kubernetes pattern, and eventually to node-local pre-staging as replica count grows.
- Implementation: ReadWriteOnce fits single-replica services; ReadWriteMany or node-local caching is required once replicas span multiple nodes.
- Versioning and registries: DVC, MLflow's Model Registry, and OCI model artifacts each give storage a version history that raw file paths cannot.
- Enterprise rollout: ownership, CI-gated promotion, scoped access to production-derived data, budget tracking, and separate runbooks for cache corruption versus registry outages turn a working pattern into a governed one.
- Common mistakes: baked-in models, wrong access modes, untrusted shared caches, overwritten datasets, over-privileged serving containers, and unmanaged retention all trace back to skipping one of the foundational questions above.
Model and dataset storage rarely fails loudly on day one — it fails quietly, months later, as an unreproducible run or a stampede during a routine rollout. Getting the mounts, access modes, and versioning right early is one of the cheapest investments in a production AI platform. Happy shipping!
Comments
Post a Comment