Imagine you are running a huge library 📚.
Every day thousands of books arrive — novels, encyclopedias, photo albums, audio recordings.
You need to store them all, find any book in seconds, and let thousands of people borrow them at the same time.
That is exactly the problem every AI team faces with data.
Training datasets, model checkpoints, experiment logs, images, audio clips — they all need a home.
Object storage is that home. And MinIO became famous as one way to build it yourself.
But here is the 2026 truth: MinIO is just one flavour of a bigger concept.
The concept — object storage with S3-compatible APIs — is what you must understand deeply.
The tool you pick (MinIO, RustFS, AWS S3, Cloudflare R2) is just a detail.
Part 1 — What Is Object Storage? (And Why Not Just Use Folders?)
You already know file storage — folders inside folders inside folders on your computer.
Like Documents → Projects → AI → Dataset → images → cat_001.jpg.
That works fine for 100 files on a laptop.
It completely breaks down for 10 billion AI training images spread across 50 servers.
The key difference:
- File storage → Data lives in a tree of nested folders. Changing the folder path breaks everything.
- Object storage → Every piece of data has a unique key (like an address). You ask for the key and get the data instantly — from any server, any location.
Think of object storage like a post office 📬 with infinite PO boxes.
Each box has a unique number (the key). You don't care which building or floor the box is on.
You just give the number and collect your parcel instantly.
Part 2 — The Three Parts of Every Object
Every piece of data stored in object storage is called an object.
Unlike a regular file, an object has three parts that travel together — always:
💡 The superpower of metadata: You can attach any custom information to an object.
For an AI training image you might store: {"class": "cat", "resolution": "1024x768", "labelled_by": "alice", "version": "v2"}.
This means you can search, filter, and organise millions of files without moving them into different folders.
Part 3 — What Is S3? The Universal Language of Object Storage
Amazon created a service called S3 (Simple Storage Service) back in 2006.
Along with it, they created an API — a set of commands — for talking to object storage.
This API became so popular that every object storage system in the world now speaks it.
💡 The S3 API is like the English language of object storage.
Just like you can speak English in London, New York, or Sydney and be understood everywhere —
you can use the S3 API to talk to AWS S3, MinIO, RustFS, Cloudflare R2, Google Cloud Storage, or any S3-compatible system.
Your code works everywhere without any changes. This is called S3 compatibility — and it is the most important concept in this entire blog post.
The four core S3 operations every MLOps engineer must know:
- PUT → Upload an object (store a file in the bucket)
- GET → Download an object (retrieve a file from the bucket)
- LIST → See all objects in a bucket (like
lsfor your storage) - DELETE → Remove an object permanently
That is it. Four simple operations. Everything else — versioning, encryption, access control, lifecycle rules — is built on top of these four.
Part 4 — What Is MinIO? The DIY Object Storage Server
Imagine you want to build your own post office inside your company building.
You don't want to rent PO boxes from Amazon (AWS S3) and pay per box forever.
You want to run it yourself, on your own servers, with your own rules.
That is exactly what MinIO was built for.
MinIO is a single binary program — you download it, run it, and instantly have your own S3-compatible server.
Your Python training scripts, your Jupyter notebooks, your ML pipelines — they all connect to it exactly like they would connect to AWS S3.
Zero code changes needed.
Part 5 — The 2026 Reality: MinIO Is One of Many
Here is the honest 2026 update every engineer needs to know:
In late 2024, MinIO moved its open-source community version to a stricter license (AGPLv3).
This created concerns for commercial use — and the community started exploring alternatives.
⚠️ The 2026 Landscape Alert — This Is Why Concepts Beat Tools:
MinIO's open-source status changed in late 2024 — its community edition moved to AGPLv3 license, which has commercial restrictions.
Multiple engineers noted: "A silent README update just ended the era of MinIO as the default open-source S3 engine."
This is exactly why this blog is about the concept (S3-compatible object storage) rather than just one tool.
The concept stays the same forever. Individual tools come and go.
✅ The Golden Rule of 2026 Object Storage:
Learn the S3 API concept and commands deeply.
Pick whichever tool fits your budget, compliance needs, and team size.
Your code will work on all of them — because they all speak S3.
Today you might use MinIO on-premise. Tomorrow you migrate to Cloudflare R2 for zero egress costs.
Your Python code changes: zero lines. That is the power of the S3 standard.
Part 6 — Key Features Every Object Storage System Has
These features exist in MinIO, AWS S3, RustFS, Cloudflare R2 — everywhere.
Learn them once and apply them everywhere.
Feature 1 — Buckets: The Top-Level Containers
A bucket is like a drawer in a filing cabinet 🗄️.
All objects must live inside a bucket. You typically create one bucket per project, team, or data category.
Feature 2 — Versioning: Never Lose Old Data
Versioning means every time you overwrite an object, the old version is kept automatically.
Think of it like Google Docs version history — every save is preserved.
- Upload
model.pt(version 1) → stored - Upload
model.ptagain (version 2) → stored alongside v1 - Something goes wrong → restore v1 instantly
This is critical for ML model management. You never accidentally lose a model checkpoint.
Feature 3 — Erasure Coding: Surviving Disk Failures
Imagine your 50 TB dataset is spread across 16 drives.
What happens when 4 drives die simultaneously? 💀
With erasure coding, object storage splits and encodes data in a way that even if several drives fail at the same time, your data is completely reconstructed automatically from the remaining drives.
Think of it like a magic jigsaw puzzle — you only need 12 of the 16 pieces to see the full picture.
Feature 4 — Lifecycle Policies: Auto-Delete Old Data
You don't want to pay storage costs for experiment runs from 3 years ago.
Lifecycle policies automatically delete or archive objects after a set time.
- Experiment logs older than 90 days → delete automatically
- Raw training data older than 1 year → move to cheap archive tier
- Model checkpoints older than v5 → remove to save space
Feature 5 — Access Control: Who Can Read What
Every bucket and object has permission rules — who can read it, who can write to it.
In MLOps you typically set:
- Training jobs → READ access to raw-data bucket, WRITE access to ml-models bucket
- Serving pods → READ-ONLY access to ml-models bucket
- Data engineers → WRITE access to raw-data bucket only
- Public users → No access at all
Part 7 — Object Storage in a Complete MLOps Pipeline
Here is where everything comes together.
Object storage is not just for holding files — it is the central nervous system connecting every part of an MLOps pipeline.
Object storage is the glue that connects every stage.
Remove it and every stage is isolated and broken.
Get it right and the entire pipeline flows automatically.
Part 8 — Choosing the Right Tool in 2026
The concept is the same everywhere. The tool choice depends on your situation.
Part 9 — When to Use Each Option (Plain English)
-
MinIO (self-hosted) →
You need full data sovereignty. Your data cannot leave your own servers (healthcare, finance, government).
You have a team that can manage infrastructure. Budget is tight and you don't want cloud bills.
Check the AGPLv3 license carefully for commercial use in 2026. -
RustFS (self-hosted, 2026 rising star) →
Same use cases as MinIO but you want Apache 2.0 license (no commercial restrictions) and faster performance.
Especially good for AI/ML data lake workloads. Community growing fast in 2025-2026. -
AWS S3 →
You are already on AWS. SageMaker, Bedrock, and EMR all integrate natively.
You need the widest tool ecosystem. You are comfortable paying for egress. -
Cloudflare R2 (2026 favourite for many teams) →
You want zero egress fees — perfect for ML inference pipelines that serve lots of data to end users.
Fully S3-compatible. Growing popularity among ML teams who move data frequently. -
DigitalOcean Spaces →
Small team, simple setup, predictable pricing. Great for indie ML projects and startups.
✅ The 2026 Practical Advice:
For learning and local development → start with MinIO or RustFS on your laptop or Minikube.
For a startup or small team on cloud → Cloudflare R2 or DigitalOcean Spaces (simple, cheap, zero egress).
For enterprise on AWS → AWS S3 with SageMaker integration.
For GDPR / data sovereignty needs → self-hosted RustFS or MinIO Enterprise.
Write your code using the boto3 Python library (the standard S3 SDK) and it runs on all of them.
❌ DON'T make these object storage mistakes:
• Don't store structured query data (rows and columns) in object storage — use a database for that
• Don't read tiny files one by one from object storage during training — batch them into larger files first
• Don't leave buckets publicly accessible — always set access policies
• Don't assume object storage is a file system — it has no real folders, only key prefixes
• Don't pick a tool without checking the license — AGPLv3 vs Apache 2.0 matters for commercial products
Quick Summary 📝
What we learned today:
- Object storage → Data stored as objects with a unique key, metadata, and the actual content. No nested folders needed.
- S3 API → The universal language of object storage. PUT, GET, LIST, DELETE. Works on every tool — AWS, MinIO, RustFS, R2.
- Bucket → The top-level container. One per project or team. All objects live inside buckets.
- MinIO → Self-hosted S3-compatible server. Your own S3 on your own hardware. Free, fast, full control.
- 2026 reality → MinIO's license changed. RustFS (Apache 2.0, 2.3x faster) is now a strong alternative.
- Core features → Versioning, erasure coding, lifecycle policies, access control — exist in every S3-compatible tool.
- MLOps role → Object storage is the central hub connecting data collection, feature engineering, training, experiment tracking, and serving.
- Tool choice → Based on cloud vs on-premise, budget, compliance, and license requirements. Code stays the same regardless.
Object storage is the foundation every MLOps engineer must understand deeply.
The tool changes. The concept stays. You now know the concept. Go build! 🪣✨
Comments
Post a Comment