Skip to main content

Posts

Showing posts with the label Hugging Face

How to Deploy a Hugging Face Space: Beginner's Guide

A Hugging Face Space is a small git repository that Hugging Face knows how to turn into a running web app — you push a few files (an app script, a dependency list, and a README with a short YAML header), and the Hub builds a container, starts it, and hands you a public URL. No servers to provision, no separate "deploy" command to memorize.  This matters because a Space is usually the fastest way to turn a model or a script into something someone else can actually click on. The friction people hit isn't the concept — it's a handful of details that changed recently or never made it into older tutorials: which SDKs are still first-class, what's free versus what now needs a paid plan, how the README's YAML block actually controls the build, and where state can (and can't) survive a restart. Get those wrong and your first Space either won't build, won't stay running, or quietly loses whatever it saved. 🧩 Same repo shape every time — the YAML...

Fine-Tune LLMs with LoRA and QLoRA Using Hugging Face PEFT

PEFT (Parameter-Efficient Fine-Tuning) is the family of techniques — and the Hugging Face library of the same name — for adapting a large pretrained model to a new task by training a small fraction of its weights instead of all of them, and LoRA (Low-Rank Adaptation) is by far its most widely used method. Instead of touching every parameter in a 7-billion or 70-billion parameter model, LoRA freezes the original weights entirely and trains a pair of tiny matrices bolted onto specific layers — often adjusting well under 1% of the total parameter count while getting most of the way to full fine-tuning quality. 🧬 This matters because it changes who gets to fine-tune a model at all. Full fine-tuning a 7B-parameter model can require well over 100GB of GPU memory once you count optimizer states and gradients — squarely hyperscaler territory. Layer in 4-bit quantization on top of LoRA (the "QLoRA" recipe), and PEFT's own documentation points to a model with 65 billion paramet...