Skip to main content

Object Storage & MinIO — The Concept That Powers Every AI Pipeline

Calculating read time…

Imagine you are running a huge library 📚.
Every day thousands of books arrive — novels, encyclopedias, photo albums, audio recordings.
You need to store them all, find any book in seconds, and let thousands of people borrow them at the same time.

That is exactly the problem every AI team faces with data.
Training datasets, model checkpoints, experiment logs, images, audio clips — they all need a home.
Object storage is that home. And MinIO became famous as one way to build it yourself.

But here is the 2026 truth: MinIO is just one flavour of a bigger concept.
The concept — object storage with S3-compatible APIs — is what you must understand deeply.
The tool you pick (MinIO, RustFS, AWS S3, Cloudflare R2) is just a detail. 




Part 1 — What Is Object Storage? (And Why Not Just Use Folders?)

You already know file storage — folders inside folders inside folders on your computer.
Like Documents → Projects → AI → Dataset → images → cat_001.jpg.

That works fine for 100 files on a laptop.
It completely breaks down for 10 billion AI training images spread across 50 servers.

📁 File Storage (Old Way) 📁 Documents 📁 Projects 📁 Music 📁 AI 📁 Dataset 🖼️ cat_001.jpg 5 hops to reach one file 😓 Breaks across multiple servers 🪣 Object Storage (Smart Way) Bucket: ai-training-data 🖼️ datasets/cats/cat_001.jpg 🤖 models/v3/weights.pt 📊 logs/run_042/metrics.csv 🎵 audio/speech_0001.wav 📄 configs/training_v2.yaml 1 unique key = instant access ✅ Works across 1000 servers equally

The key difference:

  • File storage → Data lives in a tree of nested folders. Changing the folder path breaks everything.
  • Object storage → Every piece of data has a unique key (like an address). You ask for the key and get the data instantly — from any server, any location.

Think of object storage like a post office 📬 with infinite PO boxes.
Each box has a unique number (the key). You don't care which building or floor the box is on.
You just give the number and collect your parcel instantly.

Part 2 — The Three Parts of Every Object

Every piece of data stored in object storage is called an object.
Unlike a regular file, an object has three parts that travel together — always:

🗂️ Anatomy of One Object 1. Data The actual content. Could be an image, a 10 GB model file, a CSV, or a zip archive. 📦 2. Metadata Information ABOUT the data. Size, type, creation date, owner, version, custom tags. 🏷️ 3. Unique Key The address. Like a PO box number. Use this to find the object anywhere, instantly. 🔑

💡 The superpower of metadata: You can attach any custom information to an object.
For an AI training image you might store: {"class": "cat", "resolution": "1024x768", "labelled_by": "alice", "version": "v2"}.
This means you can search, filter, and organise millions of files without moving them into different folders.

Part 3 — What Is S3? The Universal Language of Object Storage

Amazon created a service called S3 (Simple Storage Service) back in 2006.
Along with it, they created an API — a set of commands — for talking to object storage.
This API became so popular that every object storage system in the world now speaks it.

💡 The S3 API is like the English language of object storage.

Just like you can speak English in London, New York, or Sydney and be understood everywhere —
you can use the S3 API to talk to AWS S3, MinIO, RustFS, Cloudflare R2, Google Cloud Storage, or any S3-compatible system.

Your code works everywhere without any changes. This is called S3 compatibility — and it is the most important concept in this entire blog post.

The four core S3 operations every MLOps engineer must know:

  • PUT → Upload an object (store a file in the bucket)
  • GET → Download an object (retrieve a file from the bucket)
  • LIST → See all objects in a bucket (like ls for your storage)
  • DELETE → Remove an object permanently

That is it. Four simple operations. Everything else — versioning, encryption, access control, lifecycle rules — is built on top of these four.

Part 4 — What Is MinIO? The DIY Object Storage Server

Imagine you want to build your own post office inside your company building.
You don't want to rent PO boxes from Amazon (AWS S3) and pay per box forever.
You want to run it yourself, on your own servers, with your own rules.
That is exactly what MinIO was built for.

Your Apps 🤖 Training script 💻 Jupyter notebook uses S3 API ↗ S3 API calls MinIO Server 🪣 Speaks S3 language Runs on YOUR servers Your data stays on-premise No cloud bill 💰 Full control & privacy Port: 9000 stores data to Your Disks 💾💾 💾💾 HDDs, SSDs, NVMe inside your servers or Kubernetes cluster Data protected by Erasure Coding ✅ Your app uses S3 language → MinIO understands it → stores on your disks

MinIO is a single binary program — you download it, run it, and instantly have your own S3-compatible server.
Your Python training scripts, your Jupyter notebooks, your ML pipelines — they all connect to it exactly like they would connect to AWS S3.
Zero code changes needed.

Part 5 — The 2026 Reality: MinIO Is One of Many

Here is the honest 2026 update every engineer needs to know:
In late 2024, MinIO moved its open-source community version to a stricter license (AGPLv3).
This created concerns for commercial use — and the community started exploring alternatives.

⚠️ The 2026 Landscape Alert — This Is Why Concepts Beat Tools:

MinIO's open-source status changed in late 2024 — its community edition moved to AGPLv3 license, which has commercial restrictions.
Multiple engineers noted: "A silent README update just ended the era of MinIO as the default open-source S3 engine."

This is exactly why this blog is about the concept (S3-compatible object storage) rather than just one tool.
The concept stays the same forever. Individual tools come and go.

S3-Compatible Object Storage The Concept Self-Hosted 🏗️ MinIO Classic choice, AGPLv3 ⚡ RustFS 2.3x faster, Apache 2.0 🌊 SeaweedFS Blobs + files, Apache 2.0 🏰 Ceph Enterprise scale, complex Cloud-Managed ☁️ AWS S3 The original, most mature 🌐 Cloudflare R2 Zero egress fees! 🔵 Azure Blob Microsoft ecosystem 🔴 DigitalOcean Simple + affordable All tools speak the same S3 language. Learn the concept once → use any tool!

✅ The Golden Rule of 2026 Object Storage:

Learn the S3 API concept and commands deeply.
Pick whichever tool fits your budget, compliance needs, and team size.
Your code will work on all of them — because they all speak S3.

Today you might use MinIO on-premise. Tomorrow you migrate to Cloudflare R2 for zero egress costs.
Your Python code changes: zero lines. That is the power of the S3 standard.

Part 6 — Key Features Every Object Storage System Has

These features exist in MinIO, AWS S3, RustFS, Cloudflare R2 — everywhere.
Learn them once and apply them everywhere.

Feature 1 — Buckets: The Top-Level Containers

A bucket is like a drawer in a filing cabinet 🗄️.
All objects must live inside a bucket. You typically create one bucket per project, team, or data category.

Typical Bucket Layout for an MLOps Project 🪣 raw-data images/ audio/ text/ Labelled training data inputs READ by training 🪣 ml-models v1/weights.pt v2/weights.pt champion/ Trained model weights & configs WRITTEN by training 🪣 experiments run_001/ run_042/ artifacts/ MLflow / W&B artifacts store READ/WRITE always 🪣 prod-outputs predictions/ reports/ logs/ Inference results from production WRITTEN by serving

Feature 2 — Versioning: Never Lose Old Data

Versioning means every time you overwrite an object, the old version is kept automatically.
Think of it like Google Docs version history — every save is preserved.

  • Upload model.pt (version 1) → stored
  • Upload model.pt again (version 2) → stored alongside v1
  • Something goes wrong → restore v1 instantly

This is critical for ML model management. You never accidentally lose a model checkpoint.

Feature 3 — Erasure Coding: Surviving Disk Failures

Imagine your 50 TB dataset is spread across 16 drives.
What happens when 4 drives die simultaneously? 💀

With erasure coding, object storage splits and encodes data in a way that even if several drives fail at the same time, your data is completely reconstructed automatically from the remaining drives.
Think of it like a magic jigsaw puzzle — you only need 12 of the 16 pieces to see the full picture.

Erasure Coding — Data Survives Drive Failures 📄 Your File 10 GB model checkpoint split 💾 D1 💾 D2 💾 D3 💾 D4 💥 D5 DEAD 💾 D6 💾 D7 💥 D8 DEAD rebuild 📄 Recovered! 100% intact from 6 of 8 drives ✅ 2 drives failed Data survived! MinIO, RustFS, AWS S3 all do this automatically.

Feature 4 — Lifecycle Policies: Auto-Delete Old Data

You don't want to pay storage costs for experiment runs from 3 years ago.
Lifecycle policies automatically delete or archive objects after a set time.

  • Experiment logs older than 90 days → delete automatically
  • Raw training data older than 1 year → move to cheap archive tier
  • Model checkpoints older than v5 → remove to save space

Feature 5 — Access Control: Who Can Read What

Every bucket and object has permission rules — who can read it, who can write to it.
In MLOps you typically set:

  • Training jobs → READ access to raw-data bucket, WRITE access to ml-models bucket
  • Serving pods → READ-ONLY access to ml-models bucket
  • Data engineers → WRITE access to raw-data bucket only
  • Public users → No access at all

Part 7 — Object Storage in a Complete MLOps Pipeline

Here is where everything comes together.
Object storage is not just for holding files — it is the central nervous system connecting every part of an MLOps pipeline.

Object Storage (MinIO / S3 / RustFS) 🪣 Buckets everywhere 1. Data Collection Sensors, APIs, scraping → raw-data bucket 2. Feature Store Reads raw-data bucket Writes features bucket 3. Model Training Reads features bucket Writes models bucket 4. Experiment Tracker MLflow artifacts → experiments bucket 5. Model Registry Points to model version in ml-models bucket 6. Model Serving Loads model from ml-models bucket Object storage sits in the middle — every step reads and writes to it

Object storage is the glue that connects every stage.
Remove it and every stage is isolated and broken.
Get it right and the entire pipeline flows automatically.

Part 8 — Choosing the Right Tool in 2026

The concept is the same everywhere. The tool choice depends on your situation.

Where does your data live? Start here 👇 On-premise / private Self-hosted S3-compatible MinIO, RustFS, SeaweedFS, Ceph Cloud / managed Cloud Object Storage AWS S3, GCS, Azure Blob, R2 RustFS Speed Apache 2.0 MinIO Mature Check license Ceph Enterprise Complex setup AWS S3 Most mature Pricey egress R2 Zero egress 2026 favourite GCS / Azure If on Google or Microsoft cloud All options above speak S3 → your code works on every single one Pick based on: cost, compliance, team expertise, and cloud vs on-premise requirements

Part 9 — When to Use Each Option (Plain English)

  • MinIO (self-hosted) →
    You need full data sovereignty. Your data cannot leave your own servers (healthcare, finance, government).
    You have a team that can manage infrastructure. Budget is tight and you don't want cloud bills.
    Check the AGPLv3 license carefully for commercial use in 2026.
  • RustFS (self-hosted, 2026 rising star) →
    Same use cases as MinIO but you want Apache 2.0 license (no commercial restrictions) and faster performance.
    Especially good for AI/ML data lake workloads. Community growing fast in 2025-2026.
  • AWS S3 →
    You are already on AWS. SageMaker, Bedrock, and EMR all integrate natively.
    You need the widest tool ecosystem. You are comfortable paying for egress.
  • Cloudflare R2 (2026 favourite for many teams) →
    You want zero egress fees — perfect for ML inference pipelines that serve lots of data to end users.
    Fully S3-compatible. Growing popularity among ML teams who move data frequently.
  • DigitalOcean Spaces →
    Small team, simple setup, predictable pricing. Great for indie ML projects and startups.

✅ The 2026 Practical Advice:

For learning and local development → start with MinIO or RustFS on your laptop or Minikube.
For a startup or small team on cloud → Cloudflare R2 or DigitalOcean Spaces (simple, cheap, zero egress).
For enterprise on AWS → AWS S3 with SageMaker integration.
For GDPR / data sovereignty needs → self-hosted RustFS or MinIO Enterprise.

Write your code using the boto3 Python library (the standard S3 SDK) and it runs on all of them.

❌ DON'T make these object storage mistakes:

• Don't store structured query data (rows and columns) in object storage — use a database for that
• Don't read tiny files one by one from object storage during training — batch them into larger files first
• Don't leave buckets publicly accessible — always set access policies
• Don't assume object storage is a file system — it has no real folders, only key prefixes
• Don't pick a tool without checking the license — AGPLv3 vs Apache 2.0 matters for commercial products

Quick Summary 📝

What we learned today:

  • Object storage → Data stored as objects with a unique key, metadata, and the actual content. No nested folders needed.
  • S3 API → The universal language of object storage. PUT, GET, LIST, DELETE. Works on every tool — AWS, MinIO, RustFS, R2.
  • Bucket → The top-level container. One per project or team. All objects live inside buckets.
  • MinIO → Self-hosted S3-compatible server. Your own S3 on your own hardware. Free, fast, full control.
  • 2026 reality → MinIO's license changed. RustFS (Apache 2.0, 2.3x faster) is now a strong alternative.
  • Core features → Versioning, erasure coding, lifecycle policies, access control — exist in every S3-compatible tool.
  • MLOps role → Object storage is the central hub connecting data collection, feature engineering, training, experiment tracking, and serving.
  • Tool choice → Based on cloud vs on-premise, budget, compliance, and license requirements. Code stays the same regardless.

Object storage is the foundation every MLOps engineer must understand deeply.
The tool changes. The concept stays. You now know the concept. Go build! 🪣✨

Comments