Have you ever baked a cake and it came out too dry? Or too sweet? 🎂
You probably adjusted the recipe — less sugar, more water, bake for fewer minutes.
Training an AI model works exactly the same way.
The "recipe settings" for an AI model are called Hyperparameters.
And Oracle Cloud Infrastructure (OCI) gives you a powerful kitchen to bake — sorry, train — the best possible models! ☁️
🔍 What Exactly is a Hyperparameter?
When an AI model learns from data, it adjusts thousands of internal numbers (called parameters).
But before learning even starts, you must set some external control dials.
These control dials are called hyperparameters.
Think of training an AI like teaching a student to ride a bicycle. 🚲
Parameters = the student's muscle memory (learned automatically by practicing).
Hyperparameters = your instructions before practice begins:
"Practice for 5 days. Take a 10-minute break every hour. Go slow at first, then speed up."
You set these before training starts. The model cannot change them on its own.
☁️ Why Train AI Models in OCI?
You could train models on your laptop — but it would take days or even weeks!
OCI gives you access to powerful GPU machines that do in hours what your laptop can't do in months.
- 🔥 NVIDIA H100 & A10 GPUs available on-demand
- 📦 OCI Object Storage — store your datasets in the cloud safely
- 📒 OCI Notebook Sessions — Jupyter notebooks on GPU machines, browser-based
- 🤗 HuggingFace integration — the most popular AI library works natively on OCI
- 💰 Pay-as-you-go — only pay for the GPU time you actually use
- 🔁 Reproducibility — save all your hyperparameter settings and re-run anytime
OCI Data Science now supports HuggingFace TrainingArguments natively inside Notebook Sessions.
You can train BERT, GPT, LLaMA-style models — all with the hyperparameters
🗺️ The Big Picture — How Hyperparameters Fit Into Training
(learning_rate, epochs, batch_size…)
(your JSONL/CoNLL dataset from OCI Storage)
(model learns batch by batch, epoch by epoch)
(how well is the model doing? check every N steps)
(load_best_model_at_end picks the winner!)
(your trained model is live!)
Every single step above is controlled by hyperparameters.
🎛️ Hyperparameter #1 — learning_rate
What Is It?
The learning rate controls how big of a step the model takes when it's learning.
After each batch of data, the model realises it made some mistakes and tries to correct them.
The learning rate decides: "How aggressively should I fix my mistakes?"
Imagine you are driving a car and you realize you are going slightly off-road. 🚗
High learning rate = Turn the wheel sharply. Fast correction — but you might oversteer and go off the other side!
Low learning rate = Turn the wheel gently. Slow correction — but you stay smooth and reach the road safely.
Perfect learning rate = Just the right turn at the right time. ✅
📊 Visual: What Happens at Different Learning Rates?
lr = 0.1Model jumps around wildly. Loss never settles. Training fails. ❌
lr = 2e-5Model learns smoothly. Loss goes down steadily. Best results! ✅
lr = 0.000001Model learns extremely slowly. Takes forever. Wastes GPU time. ⚠️
🎯 Recommended Values
- Fine-tuning BERT/RoBERTa:
1e-5to5e-5(i.e., 0.00001 to 0.00005) - Fine-tuning LLaMA / Mistral:
1e-4to3e-4 - Training from scratch:
1e-3to1e-2 - OCI Generative AI fine-tuning: Default
2e-5is usually a great starting point
🎛️ Hyperparameter #2 — num_train_epochs
🤔 What Is It?
An epoch means: "the model has seen every single training example exactly once."
num_train_epochs is how many times the model goes through the entire dataset.
Imagine your training data is a textbook. 📚
1 epoch = You read the textbook once. You learned some things but missed a lot.
3 epochs = You read it 3 times. Now you understand it well!
100 epochs = You read it 100 times. Now you've memorised every sentence exactly.
You pass exams using only that book — but struggle with any new question. This is called overfitting! ❌
📊 What Happens at Different Epoch Counts?
🎯 Recommended Values
- Fine-tuning pre-trained models (BERT, GPT):
3to5epochs - LLM fine-tuning (LLaMA, Mistral):
1to3epochs (they learn fast!) - Small datasets (<1000 samples): Up to
10epochs may be needed - Large datasets (100k+ samples): Even
1epoch can be enough
🎛️ Hyperparameter #3 — weight_decay
🤔 What Is It?
weight_decay is a regularization technique.
It gently pushes the model's internal numbers (weights) towards zero during training.
This prevents the model from becoming too "complicated" — which leads to overfitting.
Imagine your model is a student and its "room" is full of stuff (weights/parameters). 🛏️
Without weight decay: The student keeps everything — old newspapers, broken toys, 500 pencils.
The room is cluttered and hard to find anything useful. (= overfitting on training noise)
With weight decay: Every day, a small rule says "throw away things you barely use."
Now the room stays clean and organised — only the truly important things remain! ✅
🔬 How It Works Mathematically (Simple Version)
During training, the model normally tries to minimise its mistakes (called loss).
Weight decay adds a small penalty: "Large weights = extra cost."
So the model naturally prefers simpler, smaller weights over giant complex ones.
🎯 Recommended Values
- Standard fine-tuning:
0.01(safe default) - Stronger regularisation needed:
0.1 - No regularisation:
0.0(use only if dataset is huge) - HuggingFace default:
0.0— but adding0.01is almost always better!
1.0).It will over-penalise the model and prevent it from learning anything meaningful.
Keep it small — between
0.0 and 0.1.
🎛️ Hyperparameter #4 — evaluation_strategy
🤔 What Is It?
evaluation_strategy tells the trainer: "When should we check how well the model is doing?"
Think of it like: when does the teacher give the student a test during their studies?
"no" → No tests at all during training. Only a final exam at the end.You won't know if the student was failing until it's too late! ❌
"epoch" → One test at the end of each "study session" (epoch).Good — you check progress regularly. ✅
"steps" → A quick quiz every N steps (like every 100 questions answered).Best for fine-grained monitoring — you catch problems early! ✅✅
📋 The Three Options Explained
| Value | Meaning | Best Use |
|---|---|---|
"no" |
Never evaluate during training | Quick experiments only — not recommended for real projects |
"epoch" |
Evaluate after each full epoch | Small to medium datasets — simple and effective |
"steps" |
Evaluate every N training steps | Large datasets where 1 epoch takes hours — monitor frequently |
If you set
evaluation_strategy="steps", you must also set eval_steps=500 (or any number).This tells the trainer: "Evaluate every 500 steps."
Without
eval_steps, the trainer won't know how often to check!
🎛️ Hyperparameter #5 — save_strategy
🤔 What Is It?
save_strategy tells the trainer: "When should we save a copy of the model to disk?"
This is critical — if training crashes (power cut, OCI instance stops), you don't lose all progress!
Imagine you are playing a long video game. 🎮
No save strategy = Playing for 5 hours with no save. Power goes out → start over! 😱
save_strategy="epoch" = The game auto-saves at the end of each level. 💾
save_strategy="steps" = The game saves every 10 minutes. Even better! ✅
In OCI, your "save location" is usually OCI Object Storage — so models are safe in the cloud!
📋 Options for save_strategy
"no"→ Never save. Use only for quick tests. Never in production!"epoch"→ Save at the end of each epoch. Good general choice."steps"→ Save every N steps (you setsave_steps). Best for long training runs.
Always set
save_strategy to match your evaluation_strategy.If you evaluate every epoch → save every epoch.
If you evaluate every 500 steps → save every 500 steps.
This allows
load_best_model_at_end=True to work correctly (more on that next!).
🎛️ Hyperparameter #6 — load_best_model_at_end
🤔 What Is It?
When set to True, this tells the trainer:
"At the very end of training, don't give me the last saved version — give me the BEST one!"
Imagine a cricket team plays 10 matches during training season. 🏏
They perform best in Match 7, but then get tired and play poorly in matches 8, 9, and 10.
load_best_model_at_end=False → We take the team from Match 10 (their worst recent form).
load_best_model_at_end=True → We take the team from Match 7 (their personal best)! 🏆
📊 How Does It Know Which Was "Best"?
The trainer tracks a metric — usually eval_loss (lower is better) or eval_accuracy (higher is better).
After every evaluation, it compares: "Is this checkpoint better than my current best?"
At the end, it loads whichever checkpoint won!
- Requires
evaluation_strategy≠"no"— you must evaluate to know what's "best" - Requires
save_strategyto matchevaluation_strategy— must save to be able to reload - Works with
metric_for_best_model— specify"eval_loss","eval_f1", etc.
If
save_strategy and evaluation_strategy do NOT match,setting
load_best_model_at_end=True will throw an error in HuggingFace Trainer.Always keep these two in sync!
🎛️ Hyperparameter #7 — per_device_train_batch_size
🤔 What Is It?
per_device_train_batch_size controls how many training examples the model looks at at one time.
The model doesn't learn from all examples simultaneously — it learns in small batches.
Imagine teaching a baby to eat. 👶🍽️
batch_size = 1 → Give the baby one spoon at a time. Very slow! But very precise.
batch_size = 8 → Give 8 spoons worth at a time. Faster! But a bit messier.
batch_size = 64 → Try to feed 64 spoons at once. Baby chokes! (GPU runs out of memory) 💥
The key is finding the biggest batch size that fits in your GPU memory without crashing.
🖥️ Batch Size and GPU Memory — The Connection
| Batch Size | GPU Memory Used | Training Speed | Gradient Quality |
|---|---|---|---|
1 |
Very Low | Very Slow 🐢 | Noisy |
8 |
Low | Moderate | Reasonable ✅ |
16 |
Medium | Fast ✅ | Good ✅ |
32 |
High | Faster ✅✅ | Smooth ✅✅ |
128+ |
Very High | Very Fast 🚀 | May need LR scaling ⚠️ |
🎯 Recommended Values for OCI GPU Shapes
- OCI VM.GPU.A10.1 (24GB VRAM): batch_size =
16to32for BERT-sized models - OCI BM.GPU.H100.8 (640GB total VRAM): batch_size =
64to128easily - If you get an OOM (Out of Memory) error: Halve your batch size and try again
- Pro tip: Use
gradient_accumulation_steps=4to simulate larger batches on smaller GPUs
⌨️ Putting It All Together — Complete Code Example
Now let's see all 7 hyperparameters working together in a real OCI Notebook Session!
We'll set up a complete HuggingFace TrainingArguments config.
This code sets up all the hyperparameters we learned into a single
TrainingArguments object.Think of it as filling in the "recipe settings" before you put the cake in the oven.
Once you create this object and pass it to the HuggingFace
Trainer, training starts with all your chosen settings!You will see comments next to each hyperparameter explaining what it does. 🎓
from transformers import TrainingArguments
# =====================================================
# COMPLETE HYPERPARAMETER CONFIGURATION FOR OCI
# =====================================================
training_args = TrainingArguments(
# ── WHERE TO SAVE THE MODEL ──────────────────────
output_dir="./oci_model_output",
# This is where checkpoints and the final model will be saved.
# In OCI Notebooks, you can also point this to an OCI Object Storage path.
# ── LEARNING RATE ────────────────────────────────
learning_rate=2e-5,
# How big each learning step is. 2e-5 = 0.00002
# Great default for fine-tuning BERT-style models.
# ── NUMBER OF EPOCHS ─────────────────────────────
num_train_epochs=3,
# The model will go through the entire training dataset 3 times.
# Usually 3 epochs is enough for fine-tuning.
# ── WEIGHT DECAY ─────────────────────────────────
weight_decay=0.01,
# Adds a small penalty on large weights to prevent overfitting.
# 0.01 is a safe, proven default value.
# ── EVALUATION STRATEGY ──────────────────────────
evaluation_strategy="epoch",
# Evaluate the model at the end of each epoch.
# Tells us: "How well is the model doing right now?"
# ── SAVE STRATEGY ────────────────────────────────
save_strategy="epoch",
# Save a model checkpoint at the end of each epoch.
# MUST match evaluation_strategy when load_best_model_at_end=True
# ── LOAD BEST MODEL ──────────────────────────────
load_best_model_at_end=True,
# At the end of training, automatically load the checkpoint
# that had the best evaluation score. Not the last one!
# ── BATCH SIZE ───────────────────────────────────
per_device_train_batch_size=16,
# How many training examples are processed at one time per GPU.
# Adjust based on your OCI GPU shape's available VRAM.
per_device_eval_batch_size=32,
# Batch size during evaluation — can be larger since no gradients needed.
# ── LOGGING ──────────────────────────────────────
logging_dir="./logs",
logging_steps=100,
# Log training progress to the console every 100 steps.
# ── REPRODUCIBILITY ──────────────────────────────
seed=42,
# Fixed random seed ensures you get the same results every run.
# Great for debugging and comparing experiments!
)
print("✅ TrainingArguments configured successfully!")
print(f" Learning Rate: {training_args.learning_rate}")
print(f" Epochs: {training_args.num_train_epochs}")
print(f" Weight Decay: {training_args.weight_decay}")
print(f" Eval Strategy: {training_args.evaluation_strategy}")
print(f" Save Strategy: {training_args.save_strategy}")
print(f" Load Best Model: {training_args.load_best_model_at_end}")
print(f" Train Batch Size: {training_args.per_device_train_batch_size}")
Output:
✅ TrainingArguments configured successfully! Learning Rate: 2e-05 Epochs: 3 Weight Decay: 0.01 Eval Strategy: epoch Save Strategy: epoch Load Best Model: True Train Batch Size: 16
All settings confirmed and ready. Now let's connect this to actual training! 🔥
This is the complete end-to-end training setup in an OCI Notebook Session.
It loads a pre-trained model (BERT), connects it with your dataset and hyperparameters,
then starts training on OCI's GPU.
Think of it as: "Turn on the oven, put in the cake mixture, and let it bake!"
from transformers import (
AutoTokenizer,
AutoModelForSequenceClassification,
Trainer,
TrainingArguments
)
from datasets import load_dataset
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
# ── STEP 1: Load a pre-trained model and tokenizer ────────────────
# We are using BERT — a popular, powerful language model from Google.
# "Sequence Classification" means we want it to classify text into categories.
model_name = "bert-base-uncased"
num_labels = 2 # Example: positive (1) or negative (0)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
num_labels=num_labels
)
print(f"✅ Model loaded: {model_name}")
# ── STEP 2: Load and tokenize your dataset ─────────────────────────
# We'll use a small dataset from HuggingFace for demonstration.
# In a real OCI project, replace this with your own JSONL from OCI Storage.
dataset = load_dataset("imdb", split={"train": "train[:1000]", "test": "test[:200]"})
# Note: We use only 1000 examples to keep this demo fast!
def tokenize_function(examples):
# Convert raw text into numbers (tokens) that BERT understands.
# max_length=128 means we cut off text after 128 words.
return tokenizer(
examples["text"],
padding="max_length",
truncation=True,
max_length=128
)
tokenized_dataset = dataset.map(tokenize_function, batched=True)
print("✅ Dataset tokenized!")
# ── STEP 3: Define an evaluation metric ────────────────────────────
# This function calculates how accurate our model is during evaluation.
# The Trainer calls this after every evaluation checkpoint.
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
accuracy = accuracy_score(labels, predictions)
f1 = f1_score(labels, predictions, average="weighted")
return {"accuracy": accuracy, "f1": f1}
# ── STEP 4: Combine model + data + hyperparameters in Trainer ───────
# The Trainer is like the "oven" — it takes everything and runs training.
trainer = Trainer(
model=model,
args=training_args, # All our hyperparameters from earlier!
train_dataset=tokenized_dataset["train"],
eval_dataset=tokenized_dataset["test"],
compute_metrics=compute_metrics
)
# ── STEP 5: Start Training! ─────────────────────────────────────────
print("\n🚀 Starting training on OCI GPU...")
trainer.train()
# ── STEP 6: Evaluate the final (best) model ─────────────────────────
results = trainer.evaluate()
print("\n📊 Final Evaluation Results:")
for key, value in results.items():
print(f" {key}: {value:.4f}")
# ── STEP 7: Save the best model ─────────────────────────────────────
trainer.save_model("./my_best_model")
print("\n✅ Best model saved to './my_best_model'")
print(" You can now upload this folder to OCI Object Storage!")
After training, this code uploads your saved model to OCI Object Storage.
Think of it like: "My cake is ready — now let me put it in the fridge (cloud storage) so I can use it anytime!"
This way, your model is safely stored in Oracle's cloud and can be loaded anytime, from anywhere.
import oci
import os
# Upload trained model folder to OCI Object Storage
config = oci.config.from_file()
object_storage = oci.object_storage.ObjectStorageClient(config)
namespace = object_storage.get_namespace().data
bucket_name = "my-models-bucket"
local_model_dir = "./my_best_model"
# Loop through all files in the saved model directory
for filename in os.listdir(local_model_dir):
local_path = os.path.join(local_model_dir, filename)
object_name = f"trained_models/sentiment_bert/{filename}"
with open(local_path, 'rb') as f:
object_storage.put_object(
namespace_name=namespace,
bucket_name=bucket_name,
object_name=object_name,
put_object_body=f
)
print(f" ✅ Uploaded: {filename}")
print(f"\n🎉 Model uploaded to OCI bucket: {bucket_name}")
Your trained model is now safe in OCI! You can load it anytime for inference or further fine-tuning. ☁️
📋 Hyperparameter Quick Reference Cheatsheet
| Hyperparameter | What It Controls | Safe Default | Danger Zone |
|---|---|---|---|
learning_rate |
How fast the model learns | 2e-5 |
> 0.01 or < 1e-7 |
num_train_epochs |
How many times data is seen | 3 |
> 10 (overfitting risk) |
weight_decay |
Prevents overfitting | 0.01 |
> 0.1 |
evaluation_strategy |
When to test the model | "epoch" |
"no" in production |
save_strategy |
When to save checkpoints | Match eval strategy | Mismatch with eval |
load_best_model_at_end |
Use best checkpoint, not last | True |
False (wastes results) |
per_device_train_batch_size |
Examples per GPU per step | 16 |
Too high = OOM crash |
🎁 Bonus — lr_scheduler_type (Trend), one more hyperparameter is becoming essential: the learning rate scheduler.
Instead of keeping the learning rate fixed throughout training, a scheduler changes it over time.
At the start of a race you sprint fast (high LR — learn quickly from new data). 🏃
Near the finish line you slow down carefully (low LR — fine-tune without overshooting). 🚶
A learning rate scheduler does this automatically for your model!
"linear"→ LR decreases in a straight line from start to end"cosine"→ LR follows a smooth wave curve — most popular for LLMs"constant"→ LR never changes — simple but often suboptimal"cosine_with_restarts"→ Cosine curve that restarts — great for long training runs
# Add this to your TrainingArguments: lr_scheduler_type = "cosine", # Smooth cosine decay — best practice warmup_ratio = 0.1, # Warm up LR for first 10% of training steps # Warmup: start with a very small LR and gradually increase it. # This prevents the model from making huge destructive updates at the start of training!
✅ Final DOs and DON'Ts
- ✅ Always set
evaluation_strategyandsave_strategyto the same value - ✅ Always use
load_best_model_at_end=Truefor real training jobs - ✅ Always add
weight_decay=0.01— it almost never hurts - ✅ Always start with a small learning rate (2e-5) and increase if training is too slow
- ✅ Always save your model to OCI Object Storage — not just local disk (which resets!)
- ✅ Always set
seed=42for reproducible experiments - ✅ Monitor your training with
logging_stepsto catch issues early
- ❌ Don't set learning rate above
0.001for fine-tuning — training will explode - ❌ Don't set
evaluation_strategy="no"for real projects — you'll have no visibility - ❌ Don't use
save_strategy="no"for long training runs — power cut = all work lost! - ❌ Don't set batch size so large that you get OOM (Out of Memory) errors
- ❌ Don't train for too many epochs on a small dataset — overfitting will ruin your model
- ❌ Don't ignore the
warmup_ratiofor LLM fine-tuning — cold starts damage model weights
📝 Quick Summary — What We Learned
- Hyperparameters → Control dials you set BEFORE training starts. The model cannot change these on its own.
- learning_rate → How big each learning step is. Too high = chaos. Too low = too slow. Sweet spot = 2e-5.
- num_train_epochs → How many times the model sees all the training data. Usually 3 is enough.
- weight_decay → Prevents the model from over-memorising training data. Default 0.01 works great.
- evaluation_strategy → When to test how well the model is doing. Use "epoch" or "steps".
- save_strategy → When to save checkpoints. Always match with evaluation_strategy.
- load_best_model_at_end → At the end, load the BEST checkpoint — not just the last one. Always set True!
- per_device_train_batch_size → How many examples per GPU per step. Bigger = faster, but needs more VRAM.
Comments
Post a Comment