Skip to main content

How to Save, Load, and Share a Fine-Tuned Model on the Hugging Face Hub

Calculating read time…

Saving, loading, and sharing a fine-tuned model is the moment your training run stops being a Jupyter notebook artifact and becomes something a teammate, a Space, or a production service can actually depend on. In the Hugging Face ecosystem, this isn't one button — it's a small set of decisions: what to save (a full checkpoint or a tiny adapter), how to save it safely, how to load the exact bytes you meant to load, and how to publish it so someone else can trust it. Get any one of those wrong and the "it works on my machine" problem doesn't go away — it just moves further downstream, into a Space that crashes, an endpoint that silently drifts, or a teammate who can't reproduce your eval score. 🧩

This matters more than it sounds like it should, because the Hub has quietly become the default place teams keep model weights — for their own private work as much as for open releases. Financial firms fine-tune small entity-recognition models and push the checkpoints back to a private Hub repo as part of the evaluation loop; fine-tuning platforms let you point a training job at a Hub repository and read or write checkpoints directly from it, so "saving the model" and "shipping the model" become the same commit. When a repo is your source of truth, a sloppy save, an unpinned from_pretrained() call, or a missing model card isn't a style nit — it's the difference between a reproducible pipeline and a production incident nobody can explain. ⚠️

Diagram of the save, load, share, reuse loop for a fine-tuned Hugging Face model
The save → load → share → reuse loop this post walks through, stage by stage.

🔀 Quick Comparison

Before the deep dive, here's the decision most people actually need to make first: what am I saving, and where does it need to live?

Approach What gets saved Typical size Best for
Full fine-tuned checkpoint Every parameter, via save_pretrained() Hundreds of MB to hundreds of GB Small/medium models, or when you need one self-contained artifact
PEFT / LoRA adapter Only the low-rank adapter matrices, via PEFT's save_pretrained() A few MB to tens of MB Large base models, many task variants, cheap storage/iteration
Merged adapter Base weights + adapter fused via merge_and_unload() Same as the base model Removing adapter-swap latency for a single production variant
Space demo A Gradio/Streamlit app pointed at a Hub repo N/A (app, not weights) Demos, feedback, internal review — not guaranteed uptime
Inference Endpoint Dedicated hardware serving a pinned repo revision N/A (managed service) Production traffic with latency/uptime expectations

1. 💾 Saving a Full Fine-Tuned Model

✅ Real example: Capital Fund Management (CFM), a Paris-based quantitative investment firm, fine-tuned compact entity-recognition models — including a GLiNER variant — on curated financial-news datasets as part of a documented Hugging Face case study. Instead of keeping the trained weights inside a notebook session, the resulting checkpoints were saved and made loadable through the standard Transformers save/load path so they could be evaluated, compared against zero-shot baselines, and deployed via Hugging Face Inference Endpoints. The point isn't the specific NER task — it's that "save the model" was treated as a first-class pipeline step, not an afterthought.

Kid analogy first: imagine you spent all afternoon customizing a sticker book — new stickers, notes in the margins, pages rearranged. If you want a friend to actually use your sticker book, you can't hand them a loose sticker and expect them to rebuild the rest from memory. You hand them the whole book: the stickers, the notes, and the table of contents that explains how it's organized. That's exactly what saving a full model does.

When you call model.save_pretrained("my-folder") on a Transformers model, you're not writing one file — you're writing a small, self-describing repository. A typical text model folder contains a config.json that records the architecture (hidden size, number of layers, attention heads — the "blueprint"), one or more .safetensors files holding the actual learned weights, and — if you also save the tokenizer — the vocabulary and tokenizer configuration. For large models, the weights are automatically sharded across multiple safetensors files, with a model.safetensors.index.json mapping each parameter name to the shard that holds it, so loading code knows exactly where to look.

Here's a minimal, original example of the pattern — not copied from any repo, just illustrating the mechanics:

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id, num_labels=3)

# ... your training loop happens here ...

save_dir = "./ticket-priority-classifier"
model.save_pretrained(save_dir)
tokenizer.save_pretrained(save_dir)

Without a matching tokenizer saved alongside the weights, whoever loads this model later has to guess which vocabulary and preprocessing rules it expects — and a guess that's even slightly wrong (a different special-token set, a different casing rule) will degrade the model silently rather than throw a loud error.

💡 Contrasting case: teams saving multi-billion-parameter models learn quickly that "save the model" also means managing disk space and shard count. Passing max_shard_size to save_pretrained() controls how large each safetensors shard is allowed to get — useful when your deployment target has a per-file size limit or slow disk I/O. Skipping this consideration on a 70B-parameter fine-tune (the same CFM-style pipeline, just at a much larger scale) can turn a five-minute save into a painful, disk-thrashing one.

One more wrinkle shows up the moment training happens across multiple GPUs with Hugging Face's accelerate library: the training-time model object has been wrapped by accelerator.prepare() into a distributed-training wrapper, and calling save_pretrained() directly on that wrapper saves the wrapper's extra layers along with your weights — producing a checkpoint that won't load cleanly back into the plain model class. Accelerate's own migration guide is explicit about the fix: call accelerator.unwrap_model(model) first.

unwrapped_model = accelerator.unwrap_model(model)
unwrapped_model.save_pretrained(
    "path/to/my_model_directory",
    is_main_process=accelerator.is_main_process,
    save_function=accelerator.save,
)

🎯 Use this when you have a complete, standalone model you want to hand off exactly as-is — no assumptions about what base model the recipient already has loaded, and no leftover distributed-training wrapper baked in.

2. 🔒 Safetensors, Pickle, and Why the File Format Is a Security Decision

✅ Real example: a public write-up co-authored by Hugging Face co-founder Thomas Wolf and Omar Sanseviero (who led Platform and Community at Hugging Face at the time) documents that the Hub runs a malware scan (built on the open-source ClamAV engine) over uploaded files, specifically because the default PyTorch/TensorFlow/scikit-learn serialization format, Python's pickle, can execute arbitrary code the moment it's deserialized. That's not a theoretical risk; it's the reason the safetensors format exists at all.

Kid analogy first: think of two lunchboxes. One is a normal sealed box — you open it, the sandwich is exactly what you packed. The other looks identical, but it has a hidden spring-loaded trapdoor that can be rigged to do something unexpected the instant you open the latch. A file format that lets arbitrary code run on load is the second lunchbox. safetensors is the first one: it stores only tensor data and metadata, nothing executable, so opening it can never trigger a hidden trapdoor.

This is why, when weights are available in safetensors format, from_pretrained() loads that format by default rather than a legacy pytorch_model.bin pickle file — it's both faster to load and structurally incapable of the deserialization exploit that plain pickle allows. It's also why a repository's "Safetensors" badge on the Hub isn't cosmetic; it's a signal about what happens when a stranger's file touches your Python process.

💡 Contrasting case: a malware scan and a safe file format reduce risk, but they don't replace judgment. A model with clean safetensors weights can still be trained on data you wouldn't want in production, or ship a model card that overstates its evaluation. Revisiting the CFM example: the team didn't stop at "the file loaded fine" — they ran their own held-out evaluation before trusting a checkpoint, which is the actual safety net.

🎯 Use this when you're pulling any model checkpoint you didn't train yourself — always prefer a safetensors-only repo and treat the scan as a floor, not a certification.

3. 📦 Loading Models the Right Way: AutoClasses and Revision Pinning

✅ Real example: Together AI's fine-tuning platform integrates directly with Hugging Face Hub repositories: customers point a training job at a Hub repo and the platform reads the base checkpoint and writes fine-tuning outputs back to it as part of the same pipeline. That kind of automated pipeline only works safely if the loading step resolves to a known, specific set of files every single run — not "whatever the latest commit happens to be today."

Kid analogy first: imagine checking a specific edition of a book out of the library — not "whatever copy is on the shelf today," but the exact 3rd edition with the page numbers your homework references. If the library quietly swapped in a 4th edition with renumbered chapters, your homework citations would suddenly point to the wrong pages. Loading a model by name alone is like asking for "that book" without specifying the edition.

The Hub's versioning is built on Git and Git LFS, which means every repository has a real commit history, branches, and tags. from_pretrained() accepts a revision argument that can be a branch name, a tag, or a full commit hash — and passing a commit hash is the only one of those three that can never change out from under you.

from transformers import AutoModelForSequenceClassification

# Unpinned: resolves to whatever "main" points to today
model = AutoModelForSequenceClassification.from_pretrained("acme-org/ticket-priority-classifier")

# Pinned: always resolves to these exact bytes, forever
model = AutoModelForSequenceClassification.from_pretrained(
    "acme-org/ticket-priority-classifier",
    revision="a1b2c3d4e5f6"  # a real commit hash from the repo's history
)

AutoClasses — AutoModelForCausalLM, AutoModelForSequenceClassification, AutoTokenizer, and friends — matter here too: they let your loading code stay identical while the underlying model architecture changes, because the correct model class is resolved from the repository's config.json rather than hardcoded by you.

💡 Contrasting case: gated model repositories add another loading wrinkle — some repositories require you to accept usage terms on the Hub website before your token is authorized to download the weights at all, which is common for research previews shared with a select group before a public release. If your CI pipeline hits a gated repo with a token that hasn't accepted those terms, the load fails in a way that looks like a network error rather than a permissions error unless you know to check.

🎯 Use this whenever loading code will run more than once — a training script, a CI job, or anything deployed — pin a commit hash, don't trust a moving branch.

4. 🗒️ Saving PEFT/LoRA Adapters Instead of Whole Models

✅ Real example: the widely used alignment-handbook workflow for instruction-tuning open models (documented in Hugging Face's PEFT and TRL library guides) trains LoRA adapters on top of frozen base models like Mistral, then publishes just the adapter weights to the Hub. Anyone who already has the base model downloaded can layer the adapter on top with PeftModel.from_pretrained() — a pattern PEFT's own documentation walks through explicitly, loading an adapter published under a name like alignment-handbook/zephyr-7b-sft-lora onto a separately-downloaded Mistral base.

Kid analogy first: imagine you borrowed a friend's textbook and, instead of rewriting the whole book, you added your own sticky notes with corrections and highlights in the margins. If you want to share your improvements, you don't need to photocopy and hand over the entire textbook — you just hand over your sticky notes, and anyone with the same textbook can stick them on themselves. That's what a LoRA adapter is: a small set of "sticky notes" (low-rank weight updates) layered on a much bigger, unchanged base model.

PEFT's save_pretrained() only writes the adapter's own files — typically adapter_config.json and adapter_model.safetensors — which is why an adapter for a multi-billion-parameter base model can still be only a few megabytes. Loading it back requires two steps: load the base model, then wrap it with the adapter.

from transformers import AutoModelForCausalLM
from peft import LoraConfig, get_peft_model, PeftModel

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2-0.5B")
lora_config = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"])
model = get_peft_model(base, lora_config)

# ... fine-tune model on your dataset ...

model.save_pretrained("./support-reply-adapter")   # a few MB, not several GB

# Later, or on a different machine:
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2-0.5B")
model = PeftModel.from_pretrained(base, "./support-reply-adapter")

# Optional: fuse the adapter into the base for lower inference latency
merged = model.merge_and_unload()
merged.save_pretrained("./support-reply-merged")

Note the last two lines carefully: merge_and_unload() is documented as not being an in-place operation — the returned object is a new, standard Transformers model, and you have to assign it to a variable and use that variable, or your merge quietly does nothing useful.

💡 Contrasting case: merging trades flexibility for speed. Once you merge and unload, you can no longer swap in a different adapter on the same base model at runtime the way you could with an unmerged PeftModel serving multiple LoRA "personalities." Diffusers users hit the same tradeoff with image-generation LoRAs: load_lora_weights() keeps the adapter swappable, while fusing and calling unload_lora_weights() locks in one style for faster, simpler inference.

🎯 Use adapters when you're iterating on many small variants of one big base model; merge only when you've settled on one variant to actually deploy.

5. 🌍 Sharing on the Hub: push_to_hub, Model Cards, and Licensing

✅ Real example: in the CFM case study, the fine-tuned entity-recognition variants weren't just kept internally — they were run through evaluation and comparison against baseline GLiNER and LLM-based approaches before anything moved further downstream, and the surrounding open-source models this work built on (GLiNER, SpanMarker) are themselves published on the Hub with their own model cards describing training data and limitations, which is exactly what let CFM adopt and adapt them with confidence in the first place.

Kid analogy first: the Hub is like a big shared class library shelf. You don't just toss your book onto the shelf — you write a label on the spine (the repository name), fill out a little card describing what the book is about and who it's for (the model card), and note any rules about who's allowed to borrow it (the license). A book with no label and no card might still be a great book, but nobody browsing the shelf can tell, and nobody will trust it enough to check it out.

Most classes in Transformers — models, tokenizers, configs, and processors — expose a push_to_hub() method, so publishing usually takes one extra line once you've already called save_pretrained() logic:

model.push_to_hub("acme-org/ticket-priority-classifier", private=True)
tokenizer.push_to_hub("acme-org/ticket-priority-classifier", private=True)

The same rules apply even if you never write a line of training code yourself: AutoTrain, Hugging Face's no-code fine-tuning UI, still produces an ordinary Transformers-format repository under the hood and pushes it to the Hub for you — the resulting model card typically notes it was "Trained Using AutoTrain" along with validation metrics, but it's loaded with the exact same from_pretrained() call, subject to the exact same licensing and revision-pinning considerations as anything you saved by hand.

But the upload itself is the easy part. The README.md in a model repository doubles as its model card, and a real one should say what data the model was trained on, what its intended use is, what its known failure modes or biases are, and what license governs reuse — because the license and any gating terms determine whether someone downstream is even legally allowed to build on your work. Skipping this step doesn't make the model less capable; it makes it untrustworthy to anyone who wasn't in the room when you trained it.

💡 Contrasting case: "private" is not the same as "safe to share broadly later." A repo pushed as private=True inside a personal namespace is easy to forget about; the same repo pushed under an organization's namespace, with the org's access controls, is what actually lets a team share a fine-tune without every member re-uploading their own copy. The next two sections cover exactly that distinction.

🎯 Use this whenever a second person or a second system will ever load what you just trained — even "just a teammate," even "just for now."

6. 🧭 Versioning and Reproducibility: Commits, Tags, and Xet Storage

✅ Real example: Hugging Face's own Hub documentation describes that, as of May 23, 2025, new Hub repositories default to the Xet storage backend, which splits large files into content-addressed chunks so that incremental commits — like re-uploading a checkpoint after one more epoch — only transfer the chunks that actually changed, rather than the whole multi-gigabyte file again. Any model, dataset, or Space repo you push to today is very likely already running on this system.

Kid analogy first: imagine your sticker book has a "history" feature — every time you rearrange or add stickers, it saves a snapshot you can go back to, and it's smart enough to only remember what actually changed between snapshots instead of copying the entire book each time. That's Git-based versioning plus chunk-level deduplication in one sentence.

Because Hub repos are Git repositories, every push is a commit with a real hash, and you can tag a specific commit the way you'd tag a software release — v1.0-eval-passed, for instance — so that "the model that passed our eval gate" is an unambiguous, retrievable object rather than a Slack message saying "the one from last Tuesday." Combined with revision pinning on the load side (Section 3), this turns "reproduce last month's result" from an archaeology project into a one-line change.

💡 Contrasting case: Xet's chunk-level deduplication is a storage and transfer optimization, not a substitute for meaningful commit messages or tags. A repo with forty untagged commits called "update" is technically fully versioned and practically unreproducible for a human trying to figure out which commit corresponds to the model currently serving production traffic.

🎯 Use tags for anything that crosses an approval gate — an eval pass, a production promotion — so the "official" commit is discoverable without reading commit messages one by one.

7. 🚀 From a Hub Repo to a Running Service: Spaces vs. Inference Endpoints

Diagram comparing a Hugging Face Space demo path against an Inference Endpoint production path, both starting from the same Hub repository

✅ Real example: the CFM workflow described earlier used Hugging Face Inference Endpoints specifically for the LLM-assisted labeling stage of their pipeline — a case where they needed dedicated, reliable compute behind an API rather than a best-effort community demo, because the labeling step fed directly into the training data for their production NER models.

Kid analogy first: a Space is like a science-fair booth — you built something cool, anyone walking by can try it, and it's totally fine if it's a little slow or occasionally needs a reboot between visitors. An Inference Endpoint is the school's actual vending machine — people expect it to work every single time, at any hour, without you standing next to it fixing jams.

A Gradio Space pointed at your Hub model is the fastest way to let people (or yourself) poke at a fine-tuned model interactively, and it's genuinely great for that job — gathering feedback, running a quick qualitative check, sharing a demo link with a stakeholder. What it is not built for is guaranteed uptime, predictable latency under load, or private networking, because it typically runs on shared or best-effort compute. An Inference Endpoint, by contrast, provisions dedicated hardware for your specific pinned repository revision and is billed for the time that hardware is running — which is exactly the tradeoff you want once real user traffic, or another production system, is depending on the response.

💡 Contrasting case: teams frequently start with a Space because it's free and fast to stand up, then get surprised when that same Space is still the thing customers are hitting six months later with no rate limiting, no monitoring, and no fallback. That jump — "demo that worked" to "load-bearing production dependency" — needs a deliberate decision to move to a managed endpoint, not something that happens by default because nobody revisited the architecture.

🎯 Use a Space to validate the idea and gather feedback; use a pinned-revision Inference Endpoint the moment something else depends on the answer being there.

8. 🏢 Rolling This Out at Organizational Scale

Kid analogy first: think of a school library with a real librarian, instead of an unattended shelf. The librarian keeps a catalog card for every book, decides which sections are open-access versus reference-only, and can tell you exactly which edition of a book is currently checked out to which classroom. That's the difference between a handful of personal Hub repos and a governed organizational Hub setup.

Once more than one person is fine-tuning and shipping models, the practices above need an organizational layer around them:

  1. Private organizations and access control: repos live under a shared organization namespace (e.g. acme-org/ticket-priority-classifier) rather than a single person's account, with role-based membership so departing team members don't take the only copy of a production model with them.
  2. Gated/licensed model governance: before anyone in the org can load a gated third-party base model, someone needs to formally accept its license terms — and someone needs to own tracking which licenses the org has accepted, since license terms vary from fully permissive to research-only to commercial-use restrictions.
  3. Model and dataset cards as internal documentation: a model card isn't just for the public Hub — internally, it's the fastest way for a new team member to learn what a checkpoint was trained on and what it's known to struggle with, instead of pinging the original author (who may have left).
  4. Revision-pinning as a production requirement: anything deployed — a Space, an Inference Endpoint, a batch job — should reference a specific commit hash or tag, never a floating branch, so an unrelated push to a shared repo can never silently change what's serving traffic.
  5. Cost governance for compute: Inference Endpoints are billed for the hardware-time they run, and GPU-backed training or fine-tuning jobs add another line item; an org-level view of which endpoints and jobs are active (and who owns each one) prevents forgotten dedicated GPUs from quietly running up a bill.
  6. CI/CD via Hub webhooks: a repo can trigger a webhook on new commits, which lets a promotion pipeline automatically kick off evaluation the moment a new fine-tune is pushed to a "candidate" branch or tag, rather than relying on someone remembering to run it manually.
  7. Evaluation gates before promotion: a new checkpoint should clear a held-out evaluation (using something like Hugging Face's Evaluate library, or a custom eval harness) and pass a review before its tag or commit becomes the one that production endpoints are pinned to.

🎯 Use this checklist the moment a model moves from "my experiment" to "something the team relies on" — retrofitting governance after an incident is much more painful than building it in from the first shared repo.

9. 🧪 Hands-On Lab: Save, Load, and Share Your First Model

This lab uses a tiny, disposable model and a throwaway repo, so there's no risk in experimenting. You'll need a free Hugging Face account and Python with transformers and huggingface_hub installed.

1
Log in from your terminal so uploads are authenticated: run huggingface-cli login and paste a write-access token from your Hugging Face account settings. Expect to see: a confirmation line telling you which account you're logged in as.
2
Load a small base model and give it a trivial "fine-tune" by just changing its label count (skip real training for this lab — the point is the save/load/share mechanics, not accuracy):
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "prajjwal1/bert-tiny"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id, num_labels=2)
Expect to see: a short download progress bar, then no errors.
3
Save it locally, then inspect the folder:
model.save_pretrained("./my-lab-model")
tokenizer.save_pretrained("./my-lab-model")
Expect to see: a config.json, a model.safetensors file, and tokenizer files in the folder — not a .bin file. Troubleshooting: if you see a .bin file instead, your installed safetensors package may be missing or outdated — reinstall it and re-run the save.
4
Push it to a throwaway private repo under your own account:
model.push_to_hub("your-username/my-lab-model-DELETE-ME", private=True)
tokenizer.push_to_hub("your-username/my-lab-model-DELETE-ME", private=True)
Expect to see: a repo URL printed in your terminal. Open it in a browser — you should see your files under "Files and versions" and a commit in the history tab.
5
Copy the commit hash from the repo's "Files and versions" history tab, then load the model back using that exact hash as the revision:
reloaded = AutoModelForSequenceClassification.from_pretrained(
    "your-username/my-lab-model-DELETE-ME",
    revision="<paste-the-commit-hash-here>"
)
Expect to see: no errors, and the model loads instantly from your local cache the second time you run it, since the hash never changes. Troubleshooting: if you get a 401/403 error, your login token likely doesn't have read access to a private repo under a different account than the one you're logged in as — double-check you're using your own username.

This toy loop is the exact same shape as the CFM-style production pattern from Section 1, just at a fraction of the scale: save → push → pin a revision → reload the pinned revision. Everything in Sections 4 through 8 (adapters, model cards, endpoints, governance) is this same loop with more structure wrapped around it.

10. 🚫 Common Mistakes (and Why They Bite)

  • Not pinning a revision, then wondering why results changed. Loading by repo name alone resolves to whatever the default branch currently points to. If a collaborator (or you, in the future) pushes an update, every downstream script silently starts using different weights with no error and no warning.
  • Ignoring a model's license or gated-access terms before shipping it. A model that's easy to download isn't automatically free to use commercially — some licenses restrict commercial use, require attribution, or forbid certain applications entirely. Finding this out after a product ships is a legal problem, not a technical one.
  • Trusting a community checkpoint in production without reading the model card or running your own eval. A high download count is a popularity signal, not a correctness guarantee. The CFM pattern of running your own evaluation before adoption exists precisely because a model that performs well on its authors' benchmark can still perform poorly on your specific data distribution.
  • Tokenizer/preprocessing mismatches between fine-tuning and inference. If the tokenizer used at inference time isn't the exact one saved alongside the fine-tuned weights — different special tokens, different max length, different padding side — the model receives inputs shaped differently than what it learned on, and accuracy degrades in a way that's hard to diagnose because nothing crashes.
  • Treating a Space demo as production-ready. No rate limiting means one enthusiastic user (or one scraper) can degrade the experience for everyone else, and no monitoring means you find out about an outage from a complaint instead of an alert.
  • Skipping a held-out evaluation split when fine-tuning. Without data the model never saw during training, a low training loss can hide memorization rather than genuine generalization — the classic case of a model that looks great on paper and falls apart the moment real users send inputs that don't resemble the training set.
  • Merging a LoRA adapter and forgetting the original adapter checkpoint. Once merged, the base-plus-adapter composition can't be cleanly un-merged; keeping the unmerged adapter around costs almost nothing in storage and preserves the option to swap it for a different base model later.
  • Saving a distributed-training-wrapped model without unwrapping it first. When training with Accelerate across multiple GPUs, the model object is wrapped for distributed execution; calling save_pretrained() on the wrapper instead of on accelerator.unwrap_model(model) bakes extra layers into the checkpoint, and the resulting weights won't load cleanly back into the plain model class.

❓ FAQ

Do I need to save the tokenizer every time I save a fine-tuned model?

Yes, unless the tokenizer is genuinely unchanged and already published in a repo everyone loading your model already has access to. In almost every practical case, saving the tokenizer alongside the model with tokenizer.save_pretrained() avoids the preprocessing-mismatch problem described in the Common Mistakes section.

What's the actual difference between a LoRA adapter repo and a full model repo on the Hub?

A full model repo contains everything needed to run the model standalone. An adapter repo contains only the small adapter weights (an adapter_config.json and adapter_model.safetensors) plus a reference to which base model it expects — it's not runnable on its own, and PEFT's loading code needs both the base model and the adapter repo to reconstruct the fine-tuned behavior.

Is it safe to load any random model I find on the Hub?

Safer than loading an arbitrary pickle file from an unknown source on the open internet, thanks to the Hub's malware scanning and the safetensors format's inability to execute code — but "safer" isn't "guaranteed safe for your use case." Read the model card, check the license, and run your own evaluation before trusting a checkpoint for anything that matters.

Should I use a Space or an Inference Endpoint to serve my fine-tuned model?

Start with a Space if you're gathering feedback or the audience is internal and latency-tolerant. Move to a dedicated Inference Endpoint once something else — a product feature, a paying customer, an automated pipeline — actually depends on the model responding reliably and quickly.

How do I know exactly which version of a model is currently deployed?

If your deployment loads the model by pinning a specific commit hash or a tag (rather than a branch like "main"), the answer is always visible in your deployment configuration itself — no guessing required. This is the single highest-leverage habit covered in this post.

🔗 References & Further Reading

Official/primary documentation relied on for technical accuracy:

Additional practitioner content (background reading, not a source of quoted or closely-followed text):

Hugging Face, the Hugging Face logo, Transformers, PEFT, Diffusers, and related names are trademarks of Hugging Face, Inc. and are used here descriptively to refer to their actual products. 

📝 Summary

  • Saving a full model with save_pretrained() writes a self-describing repo: config, safetensors weights, and (if you remember to) the tokenizer.
  • Safetensors exists for security — it can't execute code on load the way pickle can, which is why it's the Hub's default and why malware scanning is a floor, not a certification.
  • Loading correctly means using AutoClasses and pinning a commit hash with revision, so results don't silently change under you.
  • PEFT/LoRA adapters let you save just the "sticky notes" on a base model — a few MB instead of several GB — and merge_and_unload() fuses them when you're ready to deploy one variant.
  • Sharing on the Hub with push_to_hub() is only half the job — the model card, license, and access settings are what make a shared model trustworthy.
  • Versioning and Xet storage give you real commit history and efficient incremental uploads; tagging approved commits makes "the production model" an unambiguous, findable object.
  • Spaces vs. Inference Endpoints is a demo-versus-production decision, not a technical detail — pick based on who's actually depending on the answer.
  • Enterprise rollout wraps all of the above in access control, license governance, cost oversight, and evaluation gates before anything is promoted.
  • Common mistakes almost all trace back to one root cause: treating "it loaded" as proof that the right bytes, with the right preprocessing, under the right license, actually loaded.


Comments