How to Fine-Tune a Pretrained Hugging Face Model on a Custom Dataset — OCI Data Science
Fine-tuning a pre-trained Hugging Face model means taking a language model that already understands grammar and meaning from millions of documents, and teaching it one narrow new skill — like telling whether two sentences mean the same thing — using a small, purpose-built dataset instead of starting from zero. 🧠
This matters because almost no enterprise team can afford to pretrain a language model from scratch — that costs millions of dollars and weeks of GPU cluster time. What every real team actually does is reuse a pretrained model and fine-tune it cheaply. Get the fine-tuning, evaluation, packaging, and deployment steps wrong, and you end up with a model that looks fine in a notebook but returns garbage — or an inference endpoint that silently fails — the moment real traffic hits it in production. ⚙️
📑 In This Post
- Sample Dataset: The Sentence Pairs You'll Train On
- Beginner Walkthrough: OCI Console Setup, Step by Step
- What Semantic Similarity Actually Means
- Why You Start From a Pretrained Model, Not From Zero
- Tokenization: Turning Sentences Into Numbers
- Building and Splitting a Trustworthy Dataset
- Fine-Tuning Mechanics: Loss, Epochs, Learning Rate
- Evaluation: Proving the Model Actually Learned
- Packaging a Reusable, Deployable Artifact
- Register the Model in OCI Model Catalog
- Create a Low-Cost Model Deployment
- Test the Deployed OCI Endpoint
- Enterprise Rollout at Scale on OCI
- Common Mistakes (and Why They Happen)
- FAQ
- References & Further Reading
- Summary
🔀 Quick Comparison: Three Ways to Get a Task-Specific Model
Before going deep on any single stage, it helps to see where fine-tuning a pretrained model actually sits among the alternatives.
| Approach | Data Needed | Compute Needed | Best For |
|---|---|---|---|
| Pretraining from scratch | Billions of tokens | GPU/TPU superclusters, weeks | Foundation model providers only |
| Fine-tuning a pretrained model | Dozens to thousands of labeled examples | One CPU or GPU shape, minutes to hours | Narrow, well-defined tasks like this one |
| Prompting a frozen model (zero/few-shot) | A handful of examples, no training | Inference only | Fast prototypes, low-volume or exploratory use |
🎯 Use this when: you have a labeled dataset, even a small one, and a task narrow enough that a general-purpose model doesn't already nail it out of the box.
1️⃣ Sample Dataset: The Sentence Pairs You'll Train On
Everything in this post is anchored to one real, small, auditable dataset: 20 OCI-support-style sentence pairs, 10 labeled similar (1) and 10 labeled different (0). It's intentionally tiny so every row is inspectable by eye — that's what makes it useful for learning the workflow, even though a production system would need far more.
| Sentence 1 | Sentence 2 | Label |
|---|---|---|
| How do I reset my OCI password? | I forgot my Oracle Cloud password. How can I change it? | 1 — similar |
| My compute instance is not starting. | My OCI VM fails to boot. | 1 — similar |
| How do I open port 443? | How can I allow HTTPS traffic in the security list? | 1 — similar |
| How do I connect to my Linux VM? | How can I SSH into my compute instance? | 1 — similar |
| How do I reset my OCI password? | How do I create a VCN? | 0 — different |
| My compute instance is not starting. | Where can I see my billing report? | 0 — different |
| How do I reset my OCI password? | How do I create an object storage bucket? | 0 — different (hard negative) |
| How do I open port 443? | How do I create a budget alert? | 0 — different (hard negative) |
Notice the last two rows on purpose: both sentences mention OCI, both use a "how do I…" pattern, yet the meanings are unrelated. Those hard negatives are what keep the model honest — without them, it can pass every "obviously different" test while secretly just checking for shared keywords.
import pandas as pd
data = [
("How do I reset my OCI password?",
"I forgot my Oracle Cloud password. How can I change it?", 1),
("My compute instance is not starting.",
"My OCI VM fails to boot.", 1),
("How do I open port 443?",
"How can I allow HTTPS traffic in the security list?", 1),
("How do I reset my OCI password?",
"How do I create a VCN?", 0),
("How do I reset my OCI password?",
"How do I create an object storage bucket?", 0), # hard negative
("How do I open port 443?",
"How do I create a budget alert?", 0), # hard negative
# ... continues to 20 rows total, 10 similar / 10 different
]
df = pd.DataFrame(data, columns=["sentence1", "sentence2", "label"])
df.to_csv("oci_semantic_similarity_small.csv", index=False)
print("Total rows:", len(df))
🎯 Use this when: you need a minimal, inspectable dataset to prove a training pipeline works end-to-end before investing in a larger labeling effort.
2️⃣ Beginner Walkthrough: OCI Console Setup, Step by Step
Kid analogy: Before a kid can start a science project at school, they first need to find the right classroom, sign in at the front desk, and get their assigned table. None of that is the experiment itself — but skip it, and there's nowhere to actually do the work. These four console steps are that "find the room and get set up" stage for OCI Data Science. 🏫
OCI Data Science sits under Analytics & AI because it's the managed environment for notebook-based experimentation, model cataloging, and model deployment — doing the work here (instead of an unmanaged local laptop) keeps everything inside your organization's cloud governance, IAM policies, and cost controls from the very first click.
Step group 1: Log in to OCI
- Open cloud.oracle.com in a browser.
- Enter your Cloud Account Name or Tenancy Name.
- Sign in using the correct identity domain or your corporate SSO login.
- After login, open the top-left navigation menu.
- Go to Analytics & AI.
Logging in through the correct tenancy and identity domain matters more than it sounds — it's what puts your work under the right security policies, compartment boundaries, and cost controls from the start, rather than in some default or unmonitored scope.
Step group 2: Open Data Science Projects
- Navigate to Analytics & AI > Data Science.
- Open Projects (or Data Science Projects).
- Click Create a new project.
| Field | Value used |
|---|---|
| Project name | semantic-similarity-learning |
| Description | Beginner project for fine-tuning and deploying a CPU-friendly Hugging Face semantic similarity model. |
| Compartment | Same demo/training compartment used for the corporate OCI environment |
A Data Science project is just an organizing container — but it's the thing that keeps notebook sessions, models, and deployments traceable to one experiment instead of scattered across unrelated compartments.
Step group 3: Create a low-cost notebook session
Inside the project, create a notebook session — this is the browser-based JupyterLab environment you'll actually write and run Python in. A small CPU-only shape is enough for a DistilBERT-scale fine-tuning job; there's no need for a GPU shape (BM.GPU, A10, A100, H100) here.
| Notebook setting | Recommended value |
|---|---|
| Name | semantic-similarity-cpu-notebook |
| Shape | VM.Standard.E4.Flex |
| OCPU | 1 |
| Memory | 16 GB |
| GPU | None |
| Block storage | Lowest accepted value, typically 50 GB |
| Networking | Default networking with internet egress |
| Extra mounts | None |
Step group 4: Open JupyterLab and verify the environment
- Wait until the notebook session status becomes Active.
- Open JupyterLab from the notebook session page.
- Create a new Python 3 notebook.
- Run a basic Python test before touching any model code.
Verifying the kernel and package environment first — before downloading or training anything — avoids confusing a plain infrastructure problem with an actual machine-learning problem later on.
print("Hello from OCI Data Science Notebook")
import sys
print(sys.version)
packages = ["torch", "transformers", "datasets", "sklearn"]
for package in packages:
try:
__import__(package)
print(package, "is installed")
except ImportError:
print(package, "is NOT installed")
✅ Worked example: if all four packages print "is installed," you can move straight to loading the tokenizer and model in the next section. If any print "is NOT installed," install them in a dedicated cell before proceeding — chasing a missing-package error mid-training is far more confusing than catching it here.
🎯 Use this when: this is your first time creating an OCI Data Science project — every later section in this guide assumes this notebook session is already Active.
3️⃣ What Semantic Similarity Actually Means
Kid analogy: Imagine two kids ask their teacher completely different-sounding questions: "Can I go to the bathroom?" and "May I please be excused for a minute?" A five-year-old who only matches exact words would say these are unrelated. A kid who actually understands language knows they mean the same thing. Semantic similarity is teaching a computer to be that second kid — to compare meaning, not just spelling. 🧒
Technically, a semantic similarity model takes two pieces of text and outputs how closely related their meaning is. That can be a continuous score (embedding cosine similarity) or, as in a sentence-pair classification setup, a discrete label: 1 if the two sentences mean roughly the same thing, 0 if they don't. Real deployments use this for FAQ matching, duplicate support-ticket detection, semantic document search, and as the retrieval layer inside a RAG system.
✅ Worked example: "How do I reset my OCI password?" and "I forgot my Oracle Cloud password, how can I change it?" are lexically almost unrelated — barely any shared words — yet they are the same support request. Label: 1 (similar). This is exactly the kind of pair a keyword search would miss, but a semantic similarity model catches.
💡 Harder example (hard negative): "How do I reset my OCI password?" and "How do I create an object storage bucket?" both mention OCI and both use a "how do I…" pattern, which can superficially look similar to a shallow model. But the meaning is completely different. Label: 0. Pairs like this — same domain vocabulary, different intent — are what actually separates a model that generalizes from one that just pattern-matches on keywords.
🎯 Use this when: you need a system to recognize that differently-worded questions or tickets mean the same thing — not when you need exact-match search.
4️⃣ Why You Start From a Pretrained Model, Not From Zero
Kid analogy: Pretraining is like a kid reading thousands of books before ever answering a single question in class. By the time the teacher asks anything, the kid already understands grammar, vocabulary, cause-and-effect, and how ideas connect — they're not memorizing answers, they're learning how language itself works. Fine-tuning is a much shorter, focused lesson afterward: "Now apply everything you already know to this one specific type of question." 📚
A widely documented, verifiable example of this idea is DistilBERT, a compact encoder produced by applying knowledge distillation to BERT. Public technical reporting on DistilBERT states that it has roughly 66 million parameters, is about 40% smaller and 60% faster than bert-base-uncased, while retaining around 97% of BERT's language-understanding performance on standard benchmarks. It was pretrained on the same English Wikipedia and BookCorpus text BERT saw, using a 6-layer, 768-dimension, 12-head architecture. That pretraining is the "years of reading" — the classification head for a new task like semantic similarity starts out randomly initialized and is the only part that must learn something genuinely new during fine-tuning.
This is precisely why AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased", num_labels=2) prints a warning about some weights not being initialized — Hugging Face is telling you, correctly, that the language-understanding backbone is fully trained, but the two-class classification head bolted on top is brand new and needs the fine-tuning step to become useful.
✅ Worked example: Loading distilbert-base-uncased and running it on a test set before any fine-tuning gives a genuine baseline — often close to random guessing on a binary task, since the classification head hasn't learned anything about "similar vs different" yet. Recording that baseline (accuracy, precision, recall, F1) is what makes the after-training comparison meaningful instead of anecdotal.
🎯 Use this when: your task is narrow, your dataset is small, and the language involved (English support tickets, product text, general prose) is already well represented in the base model's original training data.
5️⃣ Tokenization: Turning Sentences Into Numbers
Kid analogy: A model can't read English any more than a calculator can read English — it only works with numbers. Tokenization is like giving every kid in class a numbered badge instead of a name tag, plus a seating chart (the attention mask) that says who's actually present versus who's an empty chair. The teacher (the model) then works entirely off badge numbers and the seating chart, never off the names themselves. 🔢
Concretely, a tokenizer splits each sentence into subword pieces from a fixed vocabulary (DistilBERT's is roughly 30,000 WordPiece tokens), converts each piece to an integer ID, and — critically for a sentence-pair task — packs both sentences into a single sequence separated by a special token, with a matching attention mask that marks real tokens versus padding. Padding and truncation to a fixed max length (128 tokens is a common, CPU-friendly choice for short support-style questions) is what lets many sentence pairs be batched together as a single tensor.
def tokenize_pair(example, tokenizer, max_length=128):
# Two sentences go in as one packed sequence with an attention mask
return tokenizer(
example["sentence1"],
example["sentence2"],
truncation=True,
padding="max_length",
max_length=max_length,
)
💡 Key warning: the tokenizer you fine-tune with must be the exact one saved alongside the fine-tuned model — same vocabulary, same special tokens, same casing convention (uncased vs cased). Swap in a mismatched tokenizer at inference time and every input_id maps to a different piece of vocabulary than the model was trained on; predictions become silently wrong with no error thrown. This is one of the most common — and most invisible — production failures in fine-tuned NLP systems.
🎯 Use this when: preparing any text for a Transformer — this step is identical whether you're building a classifier, an embedding model, or a generation model.
6️⃣ Building and Splitting a Trustworthy Dataset
Kid analogy: Imagine studying for a test using the exact same questions that will be on the exam. You'd score perfectly — and learn nothing about whether you actually understand the material. A train/test split is the teacher holding back a few questions you've never seen, specifically so your score means something. 🎓
A documented small-scale reference build for this exact task used 20 labeled sentence pairs — 10 similar, 10 different — split 80/20 into 16 training rows and 4 test rows, using stratified sampling so both classes appear in each split. That's intentionally tiny, useful for proving the mechanics end-to-end cheaply, but it also demonstrates the single biggest limitation of small datasets: with only 4 test examples, accuracy and F1 can swing wildly on the addition or removal of a single row. The documented next step for that project — expanding to roughly 100 rows with deliberate hard negatives (pairs that share vocabulary but differ in meaning) — is the correct fix, not a bigger model.
- Collect or write sentence pairs with clear ground-truth labels (1 = similar, 0 = different).
- Deliberately include hard negatives — same domain, different intent — not just obviously unrelated pairs.
- De-duplicate exact and near-duplicate rows so the model doesn't see the same signal repeated with no added information.
- Split with stratification so each class is proportionally represented in both train and test sets.
- Never let any test-set example leak into training, even accidentally through duplicate rows across the split.
✅ Worked example: "My database backup failed" / "The DB backup job did not complete successfully" (label 1) and "My compute instance is not starting" / "Where can I see my billing report" (label 0) give the model two clearly opposite signals to learn from — different phrasing, same meaning versus different phrasing, different meaning.
🎯 Use this when: designing any labeled dataset for classification — the stratify-and-hold-out pattern applies far beyond semantic similarity.
7️⃣ Fine-Tuning Mechanics: Loss, Epochs, Learning Rate
Kid analogy: Think of fine-tuning as a kid practicing free throws. Too few practice sessions (too few epochs) and the form never improves. Way too many, and the kid overcorrects for one specific hoop's rim and stops shooting well anywhere else — that's overfitting. The learning rate is how big each correction is: nudge form slightly after each miss (small learning rate) and improvement is slow but stable; swing wildly after every miss (large learning rate) and the kid never settles into good form at all. 🏀
Mechanically, each training step runs a forward pass (predict a label), computes cross-entropy loss against the true label, backpropagates the gradient, and nudges every trainable weight — both the pretrained backbone and the new classification head — a small step in the direction that reduces loss. A CPU-friendly configuration for a tiny dataset typically uses a small batch size (4 is common for a handful of training rows), a conservative learning rate around 2e-5 (a standard starting point for fine-tuning BERT-family encoders), and only 2–4 epochs, since more passes over a very small dataset mainly risks memorizing it rather than generalizing.
from transformers import TrainingArguments, Trainer
training_args = TrainingArguments(
output_dir="./similarity_model",
num_train_epochs=3,
per_device_train_batch_size=4,
learning_rate=2e-5,
weight_decay=0.01,
eval_strategy="epoch", # renamed from evaluation_strategy in
# Transformers 4.46 — the old name still
# works but is deprecated
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=test_dataset,
compute_metrics=compute_metrics,
)
trainer.train()
💡 Key warning: the argument name evaluation_strategy was deprecated in favor of eval_strategy and removed in Transformers 4.46 — a script written against an older tutorial will throw on a newer library version. The same release cycle also began deprecating the tokenizer= argument to Trainer() in favor of processing_class=. Always check your installed transformers version against current release notes before copying training code from an older guide.
🎯 Use this when: fine-tuning any encoder-based classification head on a small labeled dataset — swap the numbers, not the pattern, for larger datasets.
8️⃣ Evaluation: Proving the Model Actually Learned
Kid analogy: Grading a true/false test only on "how many did you get right" hides important information — a kid who answers "true" to everything can score 50% by luck. Precision, recall, and F1 are like asking more specific follow-up questions: of the times you said "true," how often were you actually right (precision)? Of all the actually-true answers, how many did you catch (recall)? F1 is the balanced score between the two. 📐
Running the same evaluation function before and after fine-tuning — on the same held-out test set — is what turns "the model seems to work" into a measurable, comparable claim. A documented run of this exact workflow reported that the deployed endpoint returned HTTP 200 successfully, but that some clearly-different sentence pairs were still misclassified as similar, precisely because the training set had only 20 rows total. That's not a bug in the code; it's the expected, honest result of training a model with almost no data, and it's exactly why the reported next step was to grow the dataset and add hard negatives rather than change the model architecture.
✅ Worked example: Testing "How can I restart my OCI compute instance?" against "What are the steps to reboot a cloud VM?" is a strong case for a fine-tuned model to catch as similar — no shared words beyond "OCI," but clearly the same intent. A model that gets this right after training but wrong before training is direct evidence the fine-tuning step worked.
🎯 Use this when: reporting any model change to a stakeholder — always pair the number with the baseline it's being compared against.
9️⃣ Packaging a Reusable, Deployable Artifact
Kid analogy: A brilliant answer scribbled only in a kid's own personal notebook, in their own shorthand, is useless to a substitute teacher. Packaging a model artifact is writing that same answer out clearly enough that anyone — or any serving system — can pick it up and use it correctly, with no extra context needed. 📦
A saved Hugging Face model directory needs specific files to be genuinely self-contained: config.json, model weights (model.safetensors or pytorch_model.bin), and the matching tokenizer files (tokenizer.json, tokenizer_config.json, vocab.txt). Reloading that saved folder in a fresh process — with no leftover notebook variables — is the real test of whether packaging worked. For OCI Data Science specifically, the serving contract is two functions in a score.py file: load_model(), called once when the endpoint starts, and predict(data, model), called on every request with the already-loaded model passed in.
MODEL_DIR = "./oci_semantic_similarity_finetuned"
def load_model():
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_DIR)
model.eval()
return {"tokenizer": tokenizer, "model": model}
def predict(data, model):
# Correct pattern: OCI calls load_model() once, then reuses it here.
# Never write predict(data, model=load_model()) — that reloads the
# model on every request and can trigger loading during import.
s1, s2 = data.get("sentence1"), data.get("sentence2")
if not s1 or not s2:
return {"error": "Please provide sentence1 and sentence2"}
inputs = model["tokenizer"](s1, s2, truncation=True,
padding=True, max_length=128,
return_tensors="pt")
with torch.no_grad():
probs = torch.softmax(model["model"](**inputs).logits, dim=1)
label = int(probs[0][1] >= probs[0][0])
return {"label": label, "confidence": round(probs[0][label].item(), 4)}
💡 Key warning: writing predict(data, model=load_model()) as a default argument silently loads the model at import time (or on every call, depending on how it's wired), instead of once at endpoint startup. That single misplaced default argument is a documented, real cause of failed or unnecessarily slow OCI Model Deployment endpoints.
🎯 Use this when: preparing any model for a serving system with a load-once/predict-many contract — this pattern shows up under different names across most model-serving platforms.
🔟 Register the Model in OCI Model Catalog
Kid analogy: A finished school project sitting loose on a kid's desk can get lost, mixed up with someone else's, or thrown out by accident. Handing it in and having the teacher log it in the class binder — with a name, a date, and a spot on the shelf — is what makes it official and findable later. Registering a model in the Model Catalog is that "hand it in and log it" step. 🗂️
Before uploading, it's worth checking the artifact's zip size. Console upload is the simplest path for a beginner, but — as covered in the enterprise rollout section below — it's capped at 100 MB, so confirming size first avoids a failed registration attempt partway through. Registering the model itself is what turns a local zip file into a governed OCI asset: discoverable by teammates, versioned, and ready to be attached to a live deployment.
- Go back to the OCI Console.
- Open your Data Science project > Models.
- Click Create model (or Register model).
- Choose the model artifact upload option.
- If needed, download the artifact zip from JupyterLab first, then upload it in the Console.
| Field | Value used |
|---|---|
| Model name | oci-semantic-similarity-distilbert |
| Description | Fine-tuned DistilBERT model for beginner semantic similarity classification using OCI support-style sentence pairs. |
| Artifact | oci_semantic_similarity_artifact.zip |
| Expected status | Active |
✅ Worked example: a zipped DistilBERT artifact for this task typically lands around 90 MB — comfortably under the 100 MB Console limit — so the Console upload path shown above is the right choice here. A larger encoder would need the SDK path covered later in this post.
🎯 Use this when: your fine-tuned artifact is ready and under 100 MB — for anything larger, jump to the enterprise rollout section for the SDK and Object Storage staging paths.
1️⃣1️⃣ Create a Low-Cost Model Deployment
Kid analogy: Logging a project in the class binder is one thing — putting it on display at the school fair, where visitors can actually walk up and ask it questions, is another. A model deployment is that display table: it takes the registered artifact and turns it into a live address anyone (or any app) can send a question to and get an answer back. 🎪
This is the step that starts costing money the moment it's active, so it deserves a deliberate, minimal configuration — and it should only run while you're actively testing or serving traffic. One instance is enough for functional validation; additional replicas only make sense once you need higher availability or real production traffic volume.
| Deployment setting | Recommended value |
|---|---|
| Name | oci-semantic-similarity-deployment |
| Model | oci-semantic-similarity-distilbert |
| Shape | Smallest CPU shape, such as VM.Standard.E4.Flex |
| OCPU | 1 |
| Memory | 16 GB |
| Instance count | 1 |
| GPU | None |
💡 Stop point before you click create: double-check CPU-only shape, 1 OCPU, 16 GB memory, instance count 1, and no GPU selected. After creation, wait for the deployment status to reach Active before sending any test requests — a request against a still-provisioning endpoint will simply fail, not indicate a real problem with your model or score.py.
🎯 Use this when: the registered model shows status Active in the Model Catalog and you're ready to expose it as a live HTTPS endpoint for testing.
1️⃣2️⃣ Test the Deployed OCI Endpoint
Kid analogy: A science-fair project that only ever worked at home, on the kid's own kitchen table, hasn't really proven anything yet. It has to work in front of visitors, on the actual display table, using the actual setup everyone else will see. Calling the live OCI endpoint — instead of just re-running predict() inside the notebook — is that real test in front of visitors. 🔬
Once the deployment reaches Active, copy the prediction endpoint URL from the deployment's details page. From the notebook, send a signed HTTPS request using OCI's resource principal signer — this proves the model genuinely works through OCI-managed infrastructure, not just inside the notebook's Python process where the model object never left memory.
import requests
from oci.auth.signers import get_resource_principals_signer
endpoint = "PASTE_YOUR_ENDPOINT_HERE"
signer = get_resource_principals_signer()
# Positive test: same meaning, different wording
payload = {
"sentence1": "How can I restart my OCI compute instance?",
"sentence2": "What are the steps to reboot a cloud VM?"
}
response = requests.post(endpoint, json=payload, auth=signer)
print("Status code:", response.status_code)
print("Response:", response.text)
# Negative test: same domain vocabulary, different intent (hard negative)
payload = {
"sentence1": "How do I reset my OCI password?",
"sentence2": "How do I create an object storage bucket?"
}
response = requests.post(endpoint, json=payload, auth=signer)
print("Status code:", response.status_code)
print("Response:", response.text)
✅ Worked example: a documented run of this exact test returned HTTP 200 for both calls — confirming the deployment and inference pipeline worked end-to-end — but the model still misclassified some hard-negative pairs as similar. That's expected with only 20 training rows, and it's exactly the signal that motivates the "grow the dataset with more hard negatives" recommendation covered earlier in this post, not a sign that the deployment itself is broken.
🎯 Use this when: validating any freshly created deployment endpoint — always test at least one positive and one hard-negative pair, not just an easy example.
1️⃣3️⃣ Enterprise Rollout at Scale on OCI
Kid analogy: A kid's science-fair project sitting on their kitchen table is fine for one kid. Displaying it at a school-wide fair means labeling it clearly, keeping it inside the assigned booth boundary, and having a teacher sign off before anyone else can touch it. Enterprise rollout is the same shift — from "it runs on my notebook" to "it runs inside governed boundaries anyone on the team can audit." 🏫
On OCI Data Science, that governed path runs through a Data Science project as the organizing container, a CPU-only notebook session for experimentation (a small VM.Standard.E4.Flex shape with 1 OCPU and modest memory is enough for a DistilBERT-scale model), and IAM/dynamic-group policies scoped to the compartment so training work stays isolated from unrelated tenants and resources.
Artifact registration is where beginner projects and production projects genuinely diverge. Current Oracle documentation states that Console upload to the Model Catalog is capped at 100 MB, while artifacts saved through the OCI Python SDK, the CLI, or the ADS SDK support uncompressed models up to 6 GB, with a separate large-model pathway supporting artifacts up to 400 GB for genuinely large models — via Object Storage staging rather than a direct browser upload. A 66-million-parameter DistilBERT checkpoint typically lands well under 100 MB zipped, which is why the Console path is viable for this exact project; a larger encoder or a multi-model ensemble would not be.
- Create or reuse a Data Science project as the compartment-scoped container for the work.
- Provision the smallest viable CPU notebook shape; avoid GPU shapes unless the model or dataset genuinely requires one.
- Register the trained artifact — Console for small zips, SDK/ADS with Object Storage staging for large ones.
- Create the Model Deployment endpoint at a fixed, minimal instance count for functional validation.
- Invoke the endpoint with a signed request and confirm a healthy HTTP 200 response.
- Deactivate or delete the notebook session and the deployment as soon as testing is complete.
Cost governance, observability, and compliance sit around every one of those steps rather than after them: notebook sessions and deployment endpoints bill for active compute, so leaving either running overnight is the single most common source of unnecessary spend on small learning projects. Least-privilege IAM and dynamic-group policies (scoped to data-science-family and object-family resources in a specific compartment) keep training tenants isolated from each other. Tracking the artifact ZIP alongside its checksum, the exact runtime.yaml conda environment slug, and the deployment shape gives you an auditable trail back from any production prediction to the exact code and weights that produced it — which matters the moment a model's behavior is questioned months later.
🎯 Use this when: moving any notebook experiment toward a shared, governed environment — the pattern (isolated compute → registered artifact → managed endpoint → mandatory cleanup) generalizes well beyond OCI or Hugging Face specifically.
1️⃣4️⃣ Common Mistakes (and Why They Happen)
| Mistake | Why It Happens | What Breaks |
|---|---|---|
| Skipping hard negatives | Writing only obviously-different pairs is faster to author | Model learns "shared keywords = similar" instead of real meaning |
| No before/after baseline | Feels like an extra step when you're eager to see training run | No way to prove fine-tuning actually improved anything |
| Tokenizer/model version mismatch | Reloading a saved model but pointing the tokenizer at a different checkpoint | Silent wrong predictions with no thrown error |
predict(data, model=load_model()) |
Looks convenient in a quick notebook test | Model reloads on every call, or loads too early during import |
| Trusting metrics from a 4-row test set | Small learning datasets naturally produce small test splits | A single flipped prediction swings accuracy by 25 percentage points |
| Leaving deployment endpoints active | Easy to forget after a successful test call | Ongoing compute cost with no active business owner |
| Using deprecated Trainer arguments | Copying code from an older tutorial or blog post | Script errors on newer transformers versions once the alias is fully removed |
❓ FAQ
Do I need a GPU to fine-tune DistilBERT on a small dataset?
No. For a dataset in the tens-to-low-hundreds of rows, a small CPU shape with 1 OCPU and modest memory is enough — training completes in minutes, not hours. GPUs matter once your dataset or batch size grows into the thousands-to-millions of rows.
Why does loading the pretrained model print a warning about uninitialized weights?
Because the classification head for your specific number of labels didn't exist in the original pretrained checkpoint — it's added fresh and starts randomly initialized. That warning is expected and simply confirms fine-tuning is required before that head is useful.
How small is too small for a training dataset?
There's no universal number, but a documented 20-row exercise for this exact task showed clearly unstable metrics and misclassified hard negatives. Roughly 100+ rows with deliberate hard negatives is a reasonable next milestone for a beginner sentence-pair classifier, and production systems typically need far more.
When does my model artifact need the OCI SDK instead of Console upload?
Once your zipped artifact exceeds 100 MB. Current Oracle documentation caps Console upload at 100 MB, while SDK, CLI, or ADS-based registration supports uncompressed models up to 6 GB, and a separate large-model workflow via Object Storage staging supports artifacts up to 400 GB.
Is sentence-pair classification the same thing as embedding-based similarity search?
No, they're related but distinct. Sentence-pair classification (what this guide covers) scores one pair at a time and needs both sentences at inference. Embedding-based similarity encodes each sentence independently into a vector once, then compares vectors with cosine similarity — which scales far better for searching across large document collections.
🔗 References & Further Reading
- Oracle Cloud Infrastructure Data Science — Model Catalog documentation (artifact size limits, model registration): docs.oracle.com/iaas/data-science/using/models-about.htm
- Oracle Cloud Infrastructure Data Science — The runtime.yaml file and score.py packaging rules: docs.cloud.oracle.com/iaas/data-science/using/model_runtime_yaml.htm
- Hugging Face Transformers documentation — Trainer and TrainingArguments API reference: huggingface.co/docs/transformers/main_classes/trainer
- Hugging Face model card — distilbert-base-uncased (architecture, parameter count, pretraining data): huggingface.co/distilbert/distilbert-base-uncased
- Sanh et al., "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter," arXiv:1910.01108: arxiv.org/abs/1910.01108
- PyTorch documentation — autograd and optimizer mechanics referenced in the fine-tuning section: pytorch.org/docs/stable/notes/autograd.html
All product names, trademarks, and registered trademarks (Oracle, Oracle Cloud Infrastructure, Hugging Face, PyTorch, and related marks) are property of their respective owners and are referenced here solely to describe compatible tooling and services. Content on this page is original synthesis and explanation written for this post — it is not reproduced or quoted verbatim from any of the sources above.
📝 Summary
- The whole guide is anchored to one real, tiny, inspectable 20-row sentence-pair dataset — including deliberate hard negatives.
- Beginners start with four console steps: log in to OCI, create a Data Science project, spin up a CPU-only notebook session, and verify the Python environment before touching any model code.
- Semantic similarity compares meaning, not just shared words — sentence-pair classification is a clear, measurable way to learn it.
- Starting from a pretrained model like DistilBERT means only a small classification head needs real training, not the whole language backbone.
- Tokenization packs sentence pairs into numerical tensors — and the tokenizer must always match the exact model it was trained with.
- Dataset quality — hard negatives, stratified splits, no leakage — matters more than dataset size for a first working model.
- Fine-tuning mechanics (loss, epochs, learning rate) are the same knobs regardless of task; small datasets need conservative settings.
- Before/after evaluation on a held-out set is what turns "it seems to work" into a defensible, measurable claim.
- A deployable artifact needs the full saved model folder plus a load-once/predict-many serving contract.
- Registering the artifact in the Model Catalog, creating a minimal single-instance CPU deployment, and testing it with signed positive and hard-negative requests closes the loop from notebook experiment to live endpoint.
- Enterprise rollout on OCI wraps every step in IAM boundaries, artifact-size-aware upload paths, and mandatory cost cleanup.
- Most real-world failures trace back to a handful of well-known mistakes — tokenizer mismatches, missing baselines, and forgotten running endpoints chief among them.
That's the full lifecycle, start to finish — from a handful of labeled sentence pairs to a governed, testable OCI endpoint. Go build something small, measure it honestly, and grow the dataset before you grow the model. 🚀
Comments
Post a Comment