Skip to main content

How to Deploy a Hugging Face Space: Beginner's Guide

Calculating read time…

A Hugging Face Space is a small git repository that Hugging Face knows how to turn into a running web app — you push a few files (an app script, a dependency list, and a README with a short YAML header), and the Hub builds a container, starts it, and hands you a public URL. No servers to provision, no separate "deploy" command to memorize. 

This matters because a Space is usually the fastest way to turn a model or a script into something someone else can actually click on. The friction people hit isn't the concept — it's a handful of details that changed recently or never made it into older tutorials: which SDKs are still first-class, what's free versus what now needs a paid plan, how the README's YAML block actually controls the build, and where state can (and can't) survive a restart. Get those wrong and your first Space either won't build, won't stay running, or quietly loses whatever it saved. 🧩

Diagram showing a Space git repository with README.md, app.py, and requirements.txt being built by the Spaces builder into a running container reachable at a public URL, with Gradio, Docker, and Static as the three SDK options feeding the same repo shape

Same repo shape every time — the YAML block in your README is what tells the builder which SDK to run.

🔀 Quick Comparison: Picking a Starting SDK

SDK Good for ZeroGPU? Cost to create
Gradio Python demos: sliders, chat, image/audio in and out, in minutes. Yes — the only SDK ZeroGPU currently supports. Free (up to 2 ZeroGPU Spaces) or PRO/Team for other hardware.
Docker Any language or framework — FastAPI, Node, a custom binary, or a template like Streamlit. No, not currently. PRO for personal accounts, Team/Enterprise for orgs.
Static Plain HTML/JS, in-browser ML (transformers.js, WebGPU), project pages. Not applicable — no server-side compute at all. Free for everyone, no plan required.

🎯 Use this when: you're deciding which SDK to pick on the New Space form before you've written a line of app code.

1. What Is a Space, Really? 🧭

Kid analogy: imagine handing a landlord a furnished apartment in a box — the furniture (your code), the utility list (your dependencies), and a note on the door saying what kind of apartment it is (your SDK). The landlord unpacks it, plugs everything in, and gives you the address. You never touched the wiring.

Under the hood, a Space is a normal Git repository, hosted the same way a model or dataset repo is on the Hub. That's not incidental — it's the whole mechanism. You clone it, add files, and push, exactly like any other Git remote; the Hub's build system watches for pushes and rebuilds automatically on every commit. There's no separate "deploy" button because the push is the deploy.

Every Space environment starts on the same baseline regardless of SDK: 2 CPU cores, 16GB of RAM, and 50GB of disk that doesn't survive a restart. What differs by SDK is how that environment gets populated — a Gradio Space installs your `requirements.txt` and runs `app.py`; a Docker Space builds whatever `Dockerfile` you provide; a Static Space just serves the files you push, with nothing to install at all.

💡 Worth noting: because everything lives in one repo, "your first Space" and "the Space powering a real product" are structurally the same thing — a README, some code, maybe a Dockerfile. The gap between a toy demo and something production-shaped is hardware, secrets, and storage settings, not a different platform.

🎯 Use this when: you want the mental model before touching the New Space form.

2. Choosing an SDK: Gradio, Docker, Static — and Where Streamlit Landed 🎛️

Kid analogy: picking an SDK is like picking a school lunch line — one is quick and pre-set (Gradio), one lets you bring literally anything from home in your own container (Docker), and one is just a sandwich you already made that needs no line at all (Static).

The Hub currently treats three SDKs as first-class options on the New Space form: Gradio, Docker, and Static (plain HTML/JS). Gradio is the fastest path for a Python demo — you describe inputs and outputs, and it builds the interface for you. Docker is the escape hatch: any language, any framework, any port, as long as you write the container yourself. Static skips a server entirely, which is why it's the one option that never requires compute of any kind.

Real example — a detail that trips up older tutorials: Streamlit used to be a fourth built-in SDK option, selectable directly from the New Space form. It no longer is. Hugging Face's own Streamlit Spaces documentation now opens with a deprecation notice: the built-in Streamlit SDK option is deprecated, and the recommended path is to pick Docker as the SDK and start from the official Streamlit template instead. If you're following a guide (including older ones) that shows a "Streamlit" button next to Gradio and Docker, expect it to be missing or relabeled on a current account.

💡 Harder case: Docker's flexibility comes at a cost — as of this writing, ZeroGPU (see section 5) only works with Gradio; no other SDK can use it. If your plan is "build a quick demo, run it free on a shared GPU slice," a Docker-based app (including a Streamlit one) currently can't take that path; it needs paid dedicated hardware instead.

🎯 Use this when: you're choosing an SDK and want to know which constraints follow from that choice later.

3. Hands-On: Build and Deploy a Real Space in About 10 Minutes 🧪

This is a small, disposable Gradio Space — a one-line sentiment checker — built to be thrown away afterward. The point isn't the model; it's watching a push turn into a running URL so the mechanics stop being abstract. Follow the cards in order.

1

Go to huggingface.co/new-space (create a free account first if you don't have one). Give it a name like my-first-space, choose Gradio as the SDK, leave hardware on CPU Basic, and set visibility to Public. Click Create Space.

2

On the new Space's Files tab, click Create a new file, name it requirements.txt, and paste two lines: gradio and transformers. Commit directly to the main branch.

3

Create a second file named app.py. Paste this exactly:

import gradio as gr
from transformers import pipeline

classifier = pipeline("sentiment-analysis")

def check(text):
    result = classifier(text)[0]
    return f"{result['label']} ({result['score']:.2f})"

demo = gr.Interface(
    fn=check,
    inputs=gr.Textbox(placeholder="Type a sentence..."),
    outputs="text",
    title="One-Line Sentiment Checker",
)

demo.launch()
4

Commit that file too, then open the App tab. You'll see a build log — installing dependencies, then downloading the default sentiment model. This takes a minute or two the first time.

5

✅ Checkpoint — expect to see this: the status badge at the top of the Space switches from Building to Running, and a text box labeled "One-Line Sentiment Checker" appears in the App tab. Type a sentence like "this tutorial is going smoothly" and submit — you should get back something like POSITIVE (0.99).

🔧 Troubleshooting the most common first build: if the App tab shows a red error instead of Running, open the Logs tab first. The single most common cause at this stage is a package used in app.py that isn't listed in requirements.txt — the log will say ModuleNotFoundError and name the missing package. Add it to requirements.txt and commit again; the rebuild triggers automatically.

That's the entire mechanism, end to end — nothing about a "real" Space changes this shape. A production version of this same repo would swap the toy model for a fine-tuned one, add a secret for a private model or API key (section 7), move to GPU or ZeroGPU hardware (section 5), and possibly add persistent storage — but it's still one README, one app file, and a git push that rebuilds on commit.

🎯 Use this when: you want to see the build-and-run mechanism once, concretely, before customizing anything.

4. The README Config Block: Metadata That Runs the Show 📋

Kid analogy: it's the shipping label on the box, not the box's contents. The builder reads the label first — SDK, entry file, port — before it ever looks inside at your actual code.

Every Space's README.md starts with a YAML block between two --- lines, and that block — not anything in the prose below it — is what the builder actually reads. The New Space form generates it for you, but editing it directly is how you change SDK versions, ports, or licensing after the fact:

---
title: One-Line Sentiment Checker
emoji: 🙂
colorFrom: blue
colorTo: green
sdk: gradio
sdk_version: "5.33.0"
app_file: app.py
pinned: false
license: mit
short_description: A tiny sentiment demo built for a walkthrough
---

A few fields carry more weight than they look like they should. sdk picks the runtime (gradio, docker, or static); app_file tells Gradio and static builds which file to actually run or serve; and for Docker Spaces specifically, app_port tells the platform which port your container listens on — the default it expects is 7860, and a container listening on a different port with no matching app_port value simply won't be reachable.

Real example: the Docker Streamlit template mentioned in section 2 sets sdk: docker together with app_port: 8501 in its README — 8501 being Streamlit's own default port — which is exactly why the template works without any extra Dockerfile fiddling: the YAML and the container agree on the port before the first request ever arrives.

This same block is also where a handful of less common but useful fields live: python_version to pin an interpreter, models and datasets to link related Hub repos on the Space's card, and hf_oauth: true to add a "Sign in with Hugging Face" button and get an authenticated user's identity inside the app. The full field list lives in the official configuration reference, linked in the references section below — worth a skim once, rather than memorized upfront.

🎯 Use this when: your Space builds but behaves oddly — wrong port, wrong entry file, wrong SDK version — and the fix is almost always in this block, not your app code.

5. Hardware Tiers: Free CPU, Paid GPU, and ZeroGPU ⚡

Diagram showing the hardware and plan ladder for Hugging Face Spaces: a Static Space costs nothing no matter who runs it, a Gradio or Docker Space on CPU Basic needs a PRO or Team plan, ZeroGPU is available on the Gradio SDK only with a free quota, and dedicated paid GPU hardware bills per hour until paused; persistent storage is shown as an optional add-on with Small, Medium, and Large tiers

Four different ways to get compute under a Space — each with its own cost and its own catch.

Kid analogy: ZeroGPU is like a shared arcade machine — you don't own it, you just walk up, play your turn, and step aside so the next person's turn can start. A dedicated GPU is more like renting the whole arcade for the hour, whether you're playing or not.

Every Space starts on CPU Basic — 2 vCPUs, 16GB RAM, free, no hourly cost. From the Space's Settings tab you can upgrade to a CPU Upgrade tier, or to one of several dedicated GPU tiers (small T4 instances at the low end, multi-GPU L40S or A100 configurations at the high end), all billed by the hour for as long as the Space stays running — including idle time, since nothing pauses a paid GPU Space automatically unless you configure it to sleep.

ZeroGPU is a different model entirely: instead of reserving a GPU for the life of the Space, it dynamically hands your app a GPU slice only for the duration of a decorated function call, then releases it. In code, this means importing the spaces package and wrapping any GPU-dependent function with @spaces.GPU:

import spaces
import gradio as gr

@spaces.GPU(duration=60)
def generate(prompt):
    # runs on a GPU worker only while this call is in flight
    return model(prompt)

gr.Interface(fn=generate, inputs="text", outputs="text").launch()

Two constraints matter more than the exact numbers: ZeroGPU is, per Hugging Face's own documentation, currently exclusive to the Gradio SDK, and GPU time is metered against a daily quota that resets on a rolling 24-hour window and scales with account tier — an unauthenticated visitor gets the least, a free registered account gets more, and PRO accounts get several times more than that, with the option to keep going past the quota on pre-paid, pay-as-you-go credits. Because the exact minute counts and underlying GPU model have both changed before and will likely change again, treat any specific figure (including ones you'll see elsewhere) as a snapshot, and check the current numbers on the official ZeroGPU documentation page before planning around them.

💡 Worth noting: anyone can use an existing public ZeroGPU Space for free, with no account required beyond what the Space itself asks for. The quota and plan requirements in this section only apply to creating and hosting your own ZeroGPU Space.

🎯 Use this when: you're deciding between "free but shared and time-limited" and "paid but dedicated and always-on" for a specific Space.

Kid analogy: think of it like a community pool with a free splash pad (Static Spaces) next to a lap pool that needs a membership card (Gradio/Docker on real hardware) — except members' kids get two free visits to the lap pool a month, no card required.

This is the detail most likely to surprise someone following an older tutorial: creating a Gradio or Docker Space that runs on compute now requires a paid plan — PRO for a personal account, Team or Enterprise for an organization. Static Spaces remain free for everyone, with no plan required, because they never touch compute in the first place.

There is one specific carve-out: a free personal account counts as being "in good standing" once its email is verified and the account has existed for at least a month, and an account meeting that bar can still create and host up to two Gradio Spaces running on ZeroGPU hardware. That ceiling isn't unique to free accounts, either — it simply scales up the paid ladder: a PRO personal account can host up to ten ZeroGPU Spaces, and a Team or Enterprise organization up to fifty, on top of whatever other hardware that plan already unlocks. That's the exact combination the hands-on walkthrough in section 3 used (Gradio SDK, and CPU Basic to start, upgradeable to ZeroGPU later): free accounts aren't locked out of building something real, they're specifically routed toward the SDK and hardware path that stays free.

💡 Harder case: a Docker Space, including the Streamlit-via-Docker path from section 2, does not fall under that ZeroGPU carve-out at all — Docker isn't ZeroGPU-compatible, so a free account looking to run a Docker-based app on real compute will hit the paid-plan requirement directly, with no free lane available.

🎯 Use this when: you're budgeting for a project and need to know whether the SDK you want is actually reachable on a free account.

7. Making State Stick: Secrets, Variables, and Persistent Storage 🔑

Kid analogy: the default disk is a whiteboard — great for working, wiped the moment the room resets. Persistent storage is a filing cabinet bolted to the floor; it survives the reset because it isn't part of the room being reset.

Two related but separate settings live under a Space's Settings tab. Variables and secrets lets you inject values into the app's environment without hardcoding them in your files: variables are plain and visible, meant for non-sensitive configuration; secrets are encrypted at rest, never shown again in the UI once saved, and never exposed in build logs — the right place for an API key or a private model's access token.

Persistent storage is a separate upgrade that mounts a real, durable disk at /data inside your Space, available in three paid tiers on top of the free, ephemeral default. It behaves like an ordinary drive — read and write to it like any filesystem path — and moving to a bigger tier is straightforward, while going smaller isn't: shrinking means deleting the current storage first and starting over. A practical trick from Hugging Face's own documentation: pointing the HF_HOME environment variable at a path under /data makes libraries like transformers and diffusers cache downloaded weights there, so a restart doesn't mean re-downloading a multi-gigabyte model from scratch.

If persistent storage isn't your goal — just occasional durable output, like a running log or a small results file — Hugging Face's own guidance points to a lighter-weight alternative: programmatically uploading files to a separate Hugging Face dataset repository using the `huggingface_hub` client library, which costs nothing beyond ordinary Hub storage and doesn't require the paid `/data` upgrade at all.

🎯 Use this when: your Space needs to remember something — an API key, a fine-tuned checkpoint, user-submitted data — across a restart.

8. Beyond the Demo: Visibility, Embedding, and Dev Mode 🖥️

Kid analogy: a Space's visibility setting is the lock on the front door; embedding is handing someone a window into your house instead of the key; Dev Mode is being allowed to rearrange the furniture while standing inside, instead of resubmitting a whole new floor plan every time.

Spaces support three visibility levels, set from the same Settings tab: public (anyone can view and use it), private (visible only to you or your org), and protected — a middle option for gated access. Any public Space can be dropped into another website via the "Embed this Space" option, which hands back a direct URL suited for use in an <iframe> — useful for putting a live demo on a portfolio page or inside product documentation without hosting it twice.

For iteration speed, Spaces also offer a Dev Mode, which opens a live, SSH-reachable connection into the running container so you can edit and test code directly inside it, rather than pushing a commit and waiting for a full rebuild for every small change. It's most useful once a Space has grown past the "one Python file" stage and rebuilds start taking long enough to interrupt a debugging session.

🎯 Use this when: you're sharing a Space outside the Hub itself, restricting who can see it, or trying to speed up a slow edit-rebuild-check loop.

9. Common Mistakes (and the Reasoning Behind Them) ⚠️

  • Following a tutorial that shows Streamlit as a built-in SDK option. It's deprecated as a first-class SDK; the current path is Docker plus the official Streamlit template. Screenshots from before this change will show a button that isn't there anymore.
  • Assuming any Gradio or Docker Space is free to create. Only Static costs nothing to spin up; Gradio and Docker on real compute need a PRO (personal) or Team/Enterprise (org) plan, with the sole exception of up to two ZeroGPU Gradio Spaces on a free account in good standing.
  • Expecting ZeroGPU to work inside a Docker Space. It doesn't, currently — ZeroGPU is Gradio-only, so a Docker-based app (including Streamlit-via-Docker) needs dedicated paid hardware for GPU access, not the free shared pool.
  • Writing to local disk and expecting it to survive a restart. The default 50GB disk is explicitly ephemeral. Anything that needs to persist — a database file, downloaded weights, user uploads — needs either the paid `/data` persistent storage upgrade or an explicit upload to a separate dataset repo.
  • Leaving a paid GPU Space running unattended. Unlike the free CPU tier, which sleeps after 48 hours of inactivity, dedicated paid hardware keeps billing by the hour until it's manually paused or downgraded — an idle GPU Space left running for a month is not a hypothetical cost, it's the default behavior.
  • Putting a secret in a "variable" instead of a "secret." Variables are visible in the Settings UI and to anyone with repo access; only the dedicated secrets field is encrypted and hidden after saving. An API key saved as a plain variable is exposed, not protected.
  • Mismatching `app_port` on a Docker Space. If your container's server listens on a port that doesn't match the `app_port` value in the README's YAML block (7860 is the default the platform expects), the Space builds successfully but the app itself is unreachable — a config problem that looks identical to a broken app from the outside.

❓ FAQ

Do I need Docker installed on my own machine to make a Docker Space?

No. The Hub builds your Dockerfile remotely — you only need Docker locally if you want to test the container yourself before pushing. Gradio and Static Spaces need even less: no local Python is strictly required for a Static Space, and Python 3.8+ is only needed locally if you're developing a Gradio app outside the browser-based file editor.

Can I switch a Space's SDK after creating it?

The SDK is set by the `sdk` field in the README's YAML block, and technically that field can be edited — but switching, say, Gradio to Docker usually means replacing most of the repo's contents anyway (an `app.py` doesn't help a Docker build). In practice it's simpler to create a new Space with the right SDK from the start.

Will my free Space get deleted if I stop using it?

A free CPU-tier Space sleeps after 48 hours of inactivity — it stops running, not gets deleted — and wakes back up the next time someone visits it, with a short cold-start delay. The repository and its code stay intact either way; only the running container and its ephemeral disk are affected.

What's actually different between a Space and an Inference Endpoint?

A Space is a full application — it can have a UI, custom logic, multiple routes, anything a container can run. An Inference Endpoint is dedicated, autoscaling infrastructure built specifically to serve one model behind an API, without a UI layer. A Gradio Space calling a model directly and a Space that instead calls a separate Inference Endpoint are two different, both valid, architectures.

Is ZeroGPU worth using instead of just paying for a small dedicated GPU?

It depends on traffic shape. ZeroGPU suits bursty, occasional usage — a demo that gets inference requests a few times a minute at most — because you're not paying for idle time between calls. A dedicated GPU makes more sense once traffic is steady enough that the GPU would rarely sit idle anyway, or once you need a Docker-based app, which ZeroGPU doesn't currently support.

🔗 References & Further Reading

Primary sources (relied on for every factual claim above):

Background reading only (used to sanity-check terminology and real-world usage, not quoted or closely followed):

Hugging Face, Spaces, ZeroGPU, and the Hugging Face Hub are trademarks of Hugging Face, Inc. Gradio, Docker, and Streamlit are trademarks of their respective owners. This post explains and synthesizes publicly documented behavior in my own words for educational purposes; it does not reproduce vendor documentation verbatim, and it is not affiliated with or endorsed by Hugging Face. Hardware specs, quotas, and pricing figures reflect publicly available documentation as of September 2026, change more often than most platform features, and should be re-checked against the official docs linked above before you budget around them.

📝 Summary

  • A Space is a git repository the Hub knows how to build and run — push a commit, and it rebuilds and restarts automatically, no separate deploy step.
  • Three SDKs are first-class today: Gradio, Docker, and Static. Streamlit's built-in SDK option is deprecated — use Docker plus the official Streamlit template instead.
  • The hands-on walkthrough in section 3 builds a disposable Gradio sentiment demo end to end: two files, one commit, one running URL, with a checkpoint and a troubleshooting note for the most common first build error.
  • The README's YAML block, not the prose below it, controls the SDK, entry file, and (for Docker) the port — most "it built but doesn't work" issues trace back to this block.
  • Static Spaces are always free. Gradio and Docker Spaces on real compute now require a PRO or Team/Enterprise plan, with a carve-out for up to two free ZeroGPU Gradio Spaces per personal account.
  • ZeroGPU gives free, shared, on-demand GPU access via the `@spaces.GPU` decorator, but is currently Gradio-only and quota-limited by account tier.
  • Secrets and variables are different tools — only secrets are encrypted and hidden after saving. Persistent storage at `/data` is a separate, paid, tiered upgrade for anything that must survive a restart.
  • Most real friction traces back to outdated assumptions (Streamlit as a built-in SDK), a missed paid-plan gate, writing to ephemeral disk, or a port mismatch in the YAML block.

One repo, one README block, one push — that's the whole deploy mechanism. Go build the sentiment demo, watch it turn green, then swap in whatever model you actually came here for. 🚀

Comments