A real-world Anomaly Detection system using a fine-tuned language model.
Our app will read server log messages and automatically decide:
🟢 "This log is NORMAL" or 🔴 "This log is ANOMALOUS (something is wrong!)"
Companies use this every day to catch system failures, security breaches, and unusual activity — before it causes damage.
🗺️ The Full Journey — At a Glance
Step 1: Get a Dataset (server logs from HuggingFace)
│
▼
Step 2: Clean the Data (remove garbage, fix formatting)
│
▼
Step 3: Validate the Data (make sure it is correct)
│
▼
Step 4: Fine-Tune the Model (teach AI about anomalies)
│
▼
Step 5: Evaluate the Model (test if it learned correctly)
│
▼
Step 6: Deploy on OCI (put it in the cloud for users)
│
▼
Step 7: Generate Inference (get real predictions!)
│
▼
Step 8: Monitor & Optimize (keep it healthy and fast)
What is Fine-Tuning?
Imagine you hire a very smart new employee. 🧑💼
They went to university and know lots of things — history, maths, science.
But they have never worked in a bakery before.
You spend two weeks training them specifically on how YOUR bakery works.
Which bread recipe you use. How you like the oven set. How to greet customers.
After that training, they are still smart — but now also a bakery expert!
Fine-tuning is exactly this for AI models.
We start with a smart general-purpose model (like a university graduate).
We then train it on our specific data (the bakery internship).
The result is a model that is both generally smart AND an expert in our domain.
BEFORE Fine-Tuning: AFTER Fine-Tuning:
┌───────────────────────┐ ┌───────────────────────────┐
│ General LLM │ │ Fine-Tuned LLM │
│ (knows everything │ ───────► │ (knows everything PLUS │
│ generally) │ │ is now an anomaly │
│ │ +Custom │ detection expert) 🏆 │
│ "What is an anomaly?"│ Data │ │
│ "Hmm, let me think..." │ │ "ERROR: disk I/O timeout │
└───────────────────────┘ │ → ANOMALOUS ⚠️" │
└───────────────────────────┘
📦 Step 1 — Get a Real Dataset from HuggingFace
We will use the HDFS Log Dataset — a publicly available dataset of
real server log messages, each labeled as NORMAL or ANOMALOUS.
This is used by real engineers at companies around the world.
We are using the
datasets library from HuggingFace to download
a ready-made log dataset directly from the internet.Think of it like downloading a textbook from an online library — one command and it is on your computer, ready to use.
We then peek at the first few rows to understand its shape.
# Install required libraries first (run this in your terminal once)
# pip install datasets pandas transformers torch
from datasets import load_dataset
import pandas as pd
# ── Download the log anomaly dataset from HuggingFace ────────────────
# "katanaml-org/logs-anomaly" is a curated dataset of server logs
# with labels: 0 = Normal, 1 = Anomalous
dataset = load_dataset("katanaml-org/logs-anomaly")
# ── Convert training split to a Pandas DataFrame for easy viewing ────
df = pd.DataFrame(dataset["train"])
# ── Peek at the first 5 rows ─────────────────────────────────────────
print(df.head())
print(f"\nTotal records: {len(df)}")
print(f"Columns: {df.columns.tolist()}")
print(f"\nLabel distribution:\n{df['label'].value_counts()}")
You will see output like this:
log_message label 0 INFO: Connection established from 192.168.1.1 0 1 ERROR: Disk I/O timeout after 30s on /dev/sda 1 2 INFO: User login successful for admin@corp.com 0 3 CRITICAL: Memory usage at 99% - OOM imminent 1 4 DEBUG: Cache flushed, 2048 entries cleared 0 Total records: 50000 Columns: ['log_message', 'label'] Label distribution: 0 38500 ← Normal logs 1 11500 ← Anomalous logs
🧹 Step 2 — Data Cleaning (Remove the Garbage!)
Raw data is always messy. Always.
Think of it like a box of LEGO bricks that also has dirt, broken pieces,
and random junk mixed in. 🧱
Before you build, you sort and clean the bricks. Same thing here.
Common problems in log data:
✗ Empty or null rows
✗ Duplicate entries (same log repeated 1000 times)
✗ Extremely long logs that will confuse the model
✗ Special characters or encoding errors
✗ IP addresses and timestamps that carry no learning value
Each function below solves one specific cleaning problem.
We chain them together like a car wash — the data goes through each step and comes out cleaner after every one.
At the end, we save the clean data to a new file so we never lose it.
import re
import pandas as pd
def remove_nulls(df):
"""Remove rows where log_message or label is missing."""
before = len(df)
df = df.dropna(subset=["log_message", "label"])
print(f"✓ Removed nulls: {before - len(df)} rows dropped")
return df
def remove_duplicates(df):
"""Remove exact duplicate log entries."""
before = len(df)
df = df.drop_duplicates(subset=["log_message"])
print(f"✓ Removed duplicates: {before - len(df)} rows dropped")
return df
def clean_log_text(text):
"""
Clean a single log message:
- Remove raw IP addresses (they don't help the model learn patterns)
- Remove timestamps like [2024-01-15 10:32:44]
- Remove excessive whitespace
- Trim to 512 characters max (models have input limits)
"""
# Remove IP addresses: patterns like 192.168.1.1
text = re.sub(r'\b\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}\b', '', text)
# Remove timestamps like 2024-01-15 10:32:44 or [2024-01-15]
text = re.sub(r'\d{4}-\d{2}-\d{2}[\sT]\d{2}:\d{2}:\d{2}', '', text)
# Remove multiple spaces, tabs, newlines → single space
text = re.sub(r'\s+', ' ', text).strip()
# Truncate very long logs to 512 characters
return text[:512]
def apply_text_cleaning(df):
"""Apply clean_log_text to every row in the dataframe."""
df["log_message"] = df["log_message"].apply(clean_log_text)
print(f"✓ Text cleaning applied to all {len(df)} rows")
return df
def remove_too_short(df, min_length=10):
"""Remove log messages that are too short to be meaningful."""
before = len(df)
df = df[df["log_message"].str.len() >= min_length]
print(f"✓ Removed too-short rows: {before - len(df)} rows dropped")
return df
# ── Run the full cleaning pipeline ───────────────────────────────────
print("=== Starting Data Cleaning Pipeline ===\n")
df = remove_nulls(df)
df = remove_duplicates(df)
df = apply_text_cleaning(df)
df = remove_too_short(df)
print(f"\n✅ Final clean dataset size: {len(df)} rows")
# ── Save clean data ───────────────────────────────────────────────────
df.to_csv("logs_clean.csv", index=False)
print("✅ Saved to logs_clean.csv")
✅ Step 3 — Data Validation (Is Our Data Actually Correct?)
Cleaning removes dirt. Validation checks that what remains actually makes sense.
Think of it like a quality inspector at a factory. 🏭
After the products are cleaned, the inspector checks them one more time.
We run a series of checks — like a checklist — on our cleaned data.
Each check asks one question: "Is this still true?"
If any check fails, we see a clear ❌ warning and can fix it before training.
Think of it as your data's health report card. 📋
import pandas as pd
df = pd.read_csv("logs_clean.csv")
print("=== Data Validation Report ===\n")
# ── Check 1: No nulls remain ─────────────────────────────────────────
null_count = df.isnull().sum().sum()
if null_count == 0:
print("✅ Check 1 PASSED: No null values found")
else:
print(f"❌ Check 1 FAILED: {null_count} null values still present!")
# ── Check 2: Labels are only 0 or 1 ──────────────────────────────────
valid_labels = set(df["label"].unique())
if valid_labels.issubset({0, 1}):
print("✅ Check 2 PASSED: All labels are valid (0 or 1)")
else:
print(f"❌ Check 2 FAILED: Found unexpected labels: {valid_labels}")
# ── Check 3: Class balance is reasonable (not too skewed) ────────────
label_counts = df["label"].value_counts()
ratio = label_counts[1] / label_counts[0]
if ratio >= 0.2:
print(f"✅ Check 3 PASSED: Class ratio is {ratio:.2f} (acceptable)")
else:
print(f"⚠️ Check 3 WARNING: Anomalies are very rare ({ratio:.2f} ratio)")
print(" Consider oversampling anomalies or using weighted loss during training.")
# ── Check 4: No log message is empty ─────────────────────────────────
empty_logs = df[df["log_message"].str.strip() == ""].shape[0]
if empty_logs == 0:
print("✅ Check 4 PASSED: No empty log messages")
else:
print(f"❌ Check 4 FAILED: {empty_logs} empty log messages found!")
# ── Check 5: Sufficient data for training ────────────────────────────
if len(df) >= 1000:
print(f"✅ Check 5 PASSED: Dataset has {len(df)} rows — sufficient for fine-tuning")
else:
print(f"❌ Check 5 FAILED: Only {len(df)} rows — consider collecting more data")
print("\n=== Validation Complete ===")
🧠 Step 4 — Fine-Tune the Model
4a. Which Model Do We Fine-Tune?
We will use DistilBERT — a lighter, faster version of Google's BERT model.
It is perfect for beginners because it:
✅ Is small enough to train on a single GPU
✅ Is already pre-trained on massive text data
✅ Reaches excellent accuracy after just a few hours of fine-tuning
Fine-Tuning Flow:
Raw Text (log message)
│
▼
┌──────────────────────┐
│ Tokenizer │ ← Converts text to numbers the model understands
│ (DistilBERT) │ "ERROR: disk timeout" → [101, 7704, 1024, 102]
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ DistilBERT Model │ ← The pre-trained brain we are upgrading
│ (pre-trained) │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Classification Head │ ← New layer we add: "Is this Normal or Anomalous?"
│ (we add this) │
└──────────┬───────────┘
│
▼
Output: 0 (Normal) or 1 (Anomalous)
We are doing five things here:
1. Load our cleaned CSV data and split it into train / validation sets.
2. Convert raw text into numbers (tokenization) that the model can read.
3. Load the DistilBERT model with a classification head on top.
4. Define training settings (how many rounds, learning speed, batch size).
5. Start the actual training loop — where the model learns from our data.
import pandas as pd
from sklearn.model_selection import train_test_split
from datasets import Dataset
from transformers import (
AutoTokenizer,
AutoModelForSequenceClassification,
TrainingArguments,
Trainer
)
import torch
# ── Step 1: Load cleaned data and split into train / validation ───────
df = pd.read_csv("logs_clean.csv")
# 80% for training, 20% for validation (checking how well it learned)
train_df, val_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df["label"])
# Convert to HuggingFace Dataset format
train_dataset = Dataset.from_pandas(train_df.rename(columns={"log_message": "text"}))
val_dataset = Dataset.from_pandas(val_df.rename(columns={"log_message": "text"}))
print(f"✅ Train: {len(train_dataset)} rows | Validation: {len(val_dataset)} rows")
# ── Step 2: Load the Tokenizer ────────────────────────────────────────
# The tokenizer converts text into token IDs (numbers)
# max_length=128 means we only look at the first 128 words of each log
MODEL_NAME = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
def tokenize(batch):
return tokenizer(
batch["text"],
padding="max_length", # Pad shorter sequences to 128
truncation=True, # Cut sequences longer than 128
max_length=128
)
train_dataset = train_dataset.map(tokenize, batched=True)
val_dataset = val_dataset.map(tokenize, batched=True)
# ── Step 3: Load the Model ────────────────────────────────────────────
# num_labels=2 means: two possible outputs (Normal or Anomalous)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME, num_labels=2)
print("✅ Model loaded successfully")
# ── Step 4: Define Training Settings ─────────────────────────────────
training_args = TrainingArguments(
output_dir="./anomaly_model", # Where to save checkpoints
num_train_epochs=3, # Train for 3 full rounds through data
per_device_train_batch_size=16, # Process 16 logs at a time
per_device_eval_batch_size=32,
evaluation_strategy="epoch", # Check accuracy after each epoch
save_strategy="epoch", # Save model after each epoch
load_best_model_at_end=True, # Keep the best version automatically
logging_dir="./logs", # Where to write training logs
logging_steps=50, # Print progress every 50 steps
learning_rate=2e-5, # How fast the model learns (small = careful)
warmup_steps=100, # Gently ramp up learning rate at start
weight_decay=0.01, # Prevent overfitting (memorizing instead of learning)
fp16=torch.cuda.is_available(), # Use faster float16 if GPU is available
)
# ── Step 5: Start Training! ───────────────────────────────────────────
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=val_dataset,
)
print("🚀 Starting fine-tuning... (this may take 30–90 mins on a single GPU)")
trainer.train()
print("✅ Training complete!")
# ── Save the final fine-tuned model ──────────────────────────────────
model.save_pretrained("./anomaly_model_final")
tokenizer.save_pretrained("./anomaly_model_final")
print("✅ Model saved to ./anomaly_model_final")
🖥️ 4b — Distributed Training on OCI GPU Clusters (For Big Datasets)
What if your dataset has 10 million logs instead of 50,000?
One GPU would take days or weeks. That is too slow.
The solution is distributed training — spreading the work across many GPUs.
Think of it like building a LEGO castle. 🏰
One person doing it alone takes 10 hours.
Ten friends each building one section? Done in 1 hour.
Distributed Training on OCI GPU Cluster: ┌──────────────────────────────────────────────────────────┐ │ OCI GPU Cluster (BM.GPU.H100.8) │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐│ │ │ GPU 0 │ │ GPU 1 │ │ GPU 2 │ │ GPU 3 ││ │ │ Batch 1 │ │ Batch 2 │ │ Batch 3 │ │ Batch 4 ││ │ │ (12500 │ │ (12500 │ │ (12500 │ │ (12500 ││ │ │ logs) │ │ logs) │ │ logs) │ │ logs) ││ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘│ │ └─────────────┴─────────────┴──────────────┘ │ │ │ │ │ Gradients averaged & │ │ shared across all GPUs │ │ │ │ │ All GPUs update │ │ model together ✅ │ └──────────────────────────────────────────────────────────┘
Instead of running
python train.py normally,we use
torchrun — a special launcher that starts your training script
on ALL available GPUs at the same time, automatically.--nproc_per_node=8 means: use 8 GPUs on this machine.The model automatically splits the data and merges the learning. 🤝
# ── Launch distributed training across 8 GPUs on an OCI H100 node ──── # Run this command in your OCI compute terminal (not in Python) torchrun \ --nproc_per_node=8 \ --master_addr=localhost \ --master_port=12355 \ train.py # ── In train.py, add these two lines at the very top ───────────────── # to make HuggingFace Trainer automatically use all GPUs: from transformers import TrainingArguments # TrainingArguments automatically detects and uses all available GPUs # when launched via torchrun — no extra code needed! 🎉 # ── For multi-NODE training (multiple OCI machines) ─────────────────── # Use OCI Data Science Distributed Training Jobs: # 1. Go to OCI Console → Data Science → Jobs # 2. Choose "Distributed Training" job type # 3. Select GPU shape: BM.GPU.H100.8 (8x H100 GPUs per node) # 4. Set number of nodes: e.g. 4 nodes = 32 GPUs total # 5. Upload your train.py and requirements.txt # 6. Click Run — OCI handles all the networking automatically ✅
BM.GPU.A10.4 — 4x NVIDIA A10 GPUs → Good for fine-tuning smaller models
BM.GPU.A100.8 — 8x NVIDIA A100 GPUs → Great for medium-scale training
BM.GPU.H100.8 — 8x NVIDIA H100 GPUs → Best for large-scale distributed training
Start with A10 for learning — it is much cheaper and still very powerful.
📊 Step 5 — Model Evaluation
Training is done. But does the model actually work?
We need to test it — like a driving exam after driving school. 🚗
We use our held-out test data (logs the model has NEVER seen before)
and measure how accurate it is.
We feed 100% new, unseen log messages to our trained model and compare
its answers to the correct labels.
We measure four key things: Accuracy, Precision, Recall, and F1 Score.
We also print a confusion matrix — a table that shows exactly where the model got things right and wrong.
from transformers import pipeline
from sklearn.metrics import (
classification_report,
confusion_matrix,
accuracy_score
)
import pandas as pd
# ── Load the fine-tuned model as a simple pipeline ────────────────────
# "text-classification" tells HuggingFace we want classification predictions
classifier = pipeline(
"text-classification",
model="./anomaly_model_final",
tokenizer="./anomaly_model_final"
)
# ── Load a separate test set (data the model NEVER saw during training) ─
test_df = pd.read_csv("logs_test.csv") # A separate file you kept aside
# ── Get model predictions ─────────────────────────────────────────────
# The model outputs a label like "LABEL_0" (Normal) or "LABEL_1" (Anomalous)
predictions = classifier(test_df["log_message"].tolist(), batch_size=32)
pred_labels = [int(p["label"].split("_")[1]) for p in predictions]
true_labels = test_df["label"].tolist()
# ── Print the full evaluation report ─────────────────────────────────
print("=== Model Evaluation Report ===\n")
print(f"Overall Accuracy: {accuracy_score(true_labels, pred_labels):.4f}")
print("\nDetailed Report:")
print(classification_report(
true_labels,
pred_labels,
target_names=["Normal (0)", "Anomalous (1)"]
))
print("\nConfusion Matrix:")
print("(Rows = Actual, Columns = Predicted)")
cm = confusion_matrix(true_labels, pred_labels)
print(f" Predicted Normal Predicted Anomalous")
print(f"Actual Normal: {cm[0][0]:>8} {cm[0][1]:>8}")
print(f"Actual Anomalous: {cm[1][0]:>8} {cm[1][1]:>8}")
A good result looks like this:
=== Model Evaluation Report ===
Overall Accuracy: 0.9412 ← 94.1% correct! Very good 🎉
Detailed Report:
precision recall f1-score
Normal (0) 0.96 0.97 0.96
Anomalous (1) 0.89 0.86 0.87
Confusion Matrix:
Predicted Normal Predicted Anomalous
Actual Normal: 7623 237
Actual Anomalous: 298 1842
← The model correctly caught 1842 out of 2140 anomalies ✅
Accuracy — Out of ALL logs, what % did we get right? (94% = great!)
Precision — When we say "ANOMALY!", how often are we correct? (89% = good)
Recall — Out of ALL real anomalies, how many did we catch? (86% = decent)
F1 Score — A single number that balances Precision and Recall. Higher = better.
For anomaly detection, Recall is most important.
Missing an anomaly (false negative) is worse than a false alarm!
☁️ Step 6 — Deploy on OCI (Make It Live for the World!)
6a. Package the Model into a Docker Image
We are packing our fine-tuned model, the FastAPI server code, and all Python libraries into one sealed container (Docker image).
Anyone — in any cloud, on any machine — can run this container and get the exact same anomaly detection behavior. No "works on my machine" problems. 📦
# ── Dockerfile ──────────────────────────────────────────────────────── # Start with a lightweight Python base that already has CUDA (for GPU inference) FROM python:3.11-slim # Set the working directory inside the container WORKDIR /app # Copy and install Python libraries first (smart caching trick) COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt # Copy our fine-tuned model files into the container COPY anomaly_model_final/ ./anomaly_model_final/ # Copy the FastAPI serving code COPY serve.py . # Our API will listen on port 8000 EXPOSE 8000 # Start the server when the container runs CMD ["uvicorn", "serve:app", "--host", "0.0.0.0", "--port", "8000"]
This is the FastAPI web server that wraps our fine-tuned model.
It listens for incoming log messages, runs them through the model, and returns a prediction (Normal or Anomalous) plus a confidence score.
Think of it as our model's customer service counter. 🛎️
# ── serve.py ──────────────────────────────────────────────────────────
from fastapi import FastAPI
from pydantic import BaseModel
from transformers import pipeline
app = FastAPI(title="AnomalyDetector API", version="1.0")
# Load the model once at startup (not on every request — that would be slow!)
print("⏳ Loading anomaly detection model...")
detector = pipeline(
"text-classification",
model="./anomaly_model_final",
tokenizer="./anomaly_model_final"
)
print("✅ Model ready!")
# Define what the incoming request should look like
class LogRequest(BaseModel):
log_message: str
# The main prediction endpoint
@app.post("/detect")
def detect_anomaly(request: LogRequest):
result = detector(request.log_message)[0]
label_id = int(result["label"].split("_")[1])
label_name = "ANOMALOUS ⚠️" if label_id == 1 else "NORMAL ✅"
confidence = round(result["score"] * 100, 2)
return {
"log_message": request.log_message,
"prediction": label_name,
"confidence": f"{confidence}%",
"raw_label": label_id
}
# Health check for Kubernetes probes
@app.get("/health")
def health():
return {"status": "AnomalyDetector is running 🟢"}
6b. Build, Push, and Deploy on OCI
We are doing three things in order:
1. Build — turn our Dockerfile into a runnable image on our laptop.
2. Push — upload that image to Oracle's cloud container storage (OCIR).
3. Deploy — tell OCI to download and run that image in the cloud.
After this, anyone in the world can use our anomaly detector via a URL. 🌍
# ── Step 1: Build the Docker image ───────────────────────────────────
docker build -t anomaly-detector:v1 .
# ── Step 2: Log in to Oracle Container Registry (OCIR) ───────────────
# Replace "ap-mumbai-1" with your actual OCI region
docker login ap-mumbai-1.ocir.io -u your-tenancy/your@email.com
# ── Step 3: Tag the image with the OCIR address ───────────────────────
docker tag anomaly-detector:v1 \
ap-mumbai-1.ocir.io/your-tenancy/anomaly-detector:v1
# ── Step 4: Push to OCIR ─────────────────────────────────────────────
docker push ap-mumbai-1.ocir.io/your-tenancy/anomaly-detector:v1
# ── Step 5: Deploy via OCI Container Instance (beginner-friendly) ─────
# Go to OCI Console → Developer Services → Container Instances → Create
# - Image: ap-mumbai-1.ocir.io/your-tenancy/anomaly-detector:v1
# - Shape: VM.Standard.E4.Flex (2 OCPUs, 8GB RAM) for CPU inference
# - Port: 8000
# - Click Create → wait 2-3 minutes → copy the public IP
# ── Step 6: Test your live endpoint ──────────────────────────────────
curl -X POST http://YOUR-PUBLIC-IP:8000/detect \
-H "Content-Type: application/json" \
-d '{"log_message": "CRITICAL: Memory at 99%, OOM kill triggered"}'
# Expected response:
{
"log_message": "CRITICAL: Memory at 99%, OOM kill triggered",
"prediction": "ANOMALOUS ⚠️",
"confidence": "97.43%",
"raw_label": 1
}
Anyone who knows your IP address can now send log messages and get anomaly predictions back.
Next step: add an OCI API Gateway in front to make the URL clean and add security.
⚡ Step 7 — Generate Inference at Scale
Your endpoint works for one request at a time.
But what if 1000 log messages arrive every second? 🏃🏃🏃
We need to handle that gracefully. Here are two techniques.
Technique A — Batch Inference (Process Many Logs at Once)
Instead of sending one log at a time (slow — like ordering pizza one slice at a time),
we send 32 logs at once in a single API call (fast — like ordering the whole pizza). 🍕
The model processes all 32 in parallel and returns all answers together.
import requests
import json
# A batch of 5 log messages to classify at once
log_batch = [
"INFO: User login successful for admin",
"CRITICAL: Disk full on /dev/sda — write failed",
"DEBUG: Cache hit ratio 94.3%",
"ERROR: SSL handshake failed after 3 retries",
"INFO: Scheduled backup completed in 2.3s"
]
# Send all 5 as a single request using the /batch endpoint
response = requests.post(
"http://YOUR-PUBLIC-IP:8000/batch-detect",
json={"logs": log_batch}
)
results = response.json()["results"]
for log, result in zip(log_batch, results):
print(f"📋 Log: {log[:50]}...")
print(f" → {result['prediction']} ({result['confidence']})\n")
Technique B — Async Inference with OCI Queue (For Very High Volume)
High Volume Async Architecture:
1000 logs/second arrive
│
▼
┌──────────────────┐
│ OCI Queue │ ← Holds logs in a line (like a ticket queue)
│ (buffer) │ Nothing gets lost even if model is busy
└──────┬───────────┘
│ Pulls batches of 32
▼
┌──────────────────┐
│ AnomalyDetector │ ← Processes at its own pace
│ (Docker on OKE) │ Multiple replicas handle load
└──────┬───────────┘
│
▼
┌──────────────────┐
│ OCI Object │ ← Results stored here for downstream systems
│ Storage │
└──────────────────┘
📈 Step 8 — Monitor Training Progress and Optimize Resources
8a. Track Training with OCI Logging and MLflow
MLflow is like a notebook that automatically records every experiment you run.
Every training run gets its own page: which settings you used, how accurate it got, how long it took, and which model file was produced.
You can then compare runs side by side to find what works best. 🔬
import mlflow
import mlflow.transformers
# ── Start an MLflow experiment ────────────────────────────────────────
mlflow.set_experiment("AnomalyDetector-FineTuning")
with mlflow.start_run(run_name="distilbert-v1"):
# ── Log all training hyperparameters ──────────────────────────────
mlflow.log_params({
"model_name": "distilbert-base-uncased",
"epochs": 3,
"batch_size": 16,
"learning_rate": 2e-5,
"max_length": 128,
"dataset_size": len(train_df)
})
# ── Log evaluation metrics after training ─────────────────────────
mlflow.log_metrics({
"accuracy": 0.9412,
"f1_anomalous": 0.87,
"precision_anomalous": 0.89,
"recall_anomalous": 0.86
})
# ── Save the model in MLflow format ──────────────────────────────
mlflow.transformers.log_model(
transformers_model={
"model": model,
"tokenizer": tokenizer
},
artifact_path="anomaly_model"
)
print("✅ Experiment logged to MLflow!")
# ── View results: run this in terminal, then open http://localhost:5000
# mlflow ui
8b. Watch GPU Usage in Real Time
During training, you want to make sure your GPU is being used efficiently.
If GPU usage is below 50%, your training is slow and you are wasting money.
These commands help you see exactly what is happening inside your GPU right now.
# ── Check GPU usage every 2 seconds (run in a separate terminal) ────── watch -n 2 nvidia-smi # ── What to look for: ───────────────────────────────────────────────── # GPU-Util: should be 85–100% during training → good use of GPU # Memory-Usage: should be 70–95% → filling GPU memory = efficient # Temperature: should stay below 85°C → safe # ── Sample good output: ─────────────────────────────────────────────── # GPU Name | GPU-Util | Memory-Usage # GPU 0 H100 80GB | 94% | 68000MiB / 81920MiB ← Great! 🟢 # ── If GPU-Util is below 50%: ───────────────────────────────────────── # → Increase batch_size (process more logs at a time) # → Use DataLoader with num_workers=4 (load data faster) # → Enable fp16=True (use faster 16-bit math)
8c. OCI Monitoring Dashboard
OCI Console → Observability & Management → Monitoring ┌─────────────────────────────────────────────────────────┐ │ AnomalyDetector API — Live Metrics Dashboard │ ├────────────────────┬────────────────────────────────────┤ │ Requests / min │ ████████████░░░ 847 req/min │ ├────────────────────┼────────────────────────────────────┤ │ Avg Latency │ ██░░░░░░░░░░░░░ 143ms │ ├────────────────────┼────────────────────────────────────┤ │ Error Rate │ ░░░░░░░░░░░░░░░ 0.2% ✅ │ ├────────────────────┼────────────────────────────────────┤ │ CPU Usage │ ████████░░░░░░░ 62% │ ├────────────────────┼────────────────────────────────────┤ │ Memory Usage │ ██████████░░░░░ 71% │ └────────────────────┴────────────────────────────────────┘ Set Alarms: → If error rate > 5% → send email to team → If latency > 500ms → auto-scale up replicas → If memory > 90% → alert on-call engineer
🔹 Use fp16=True — cuts memory usage by 50%, speeds up training by 30–40%
🔹 Increase batch_size as high as your GPU memory allows — better GPU utilization
🔹 Use gradient_accumulation_steps=4 to simulate larger batches on small GPUs
🔹 Use OCI Spot Instances for training — up to 70% cheaper than on-demand shapes
🔹 Always stop your training GPU instance after training — idle GPUs still cost money!
❌ Never train on your FULL dataset without validating a small sample first.
❌ Never skip data cleaning — garbage in = garbage model.
❌ Never use a learning rate above 5e-5 for fine-tuning — it destroys the pre-trained knowledge.
❌ Never leave OCI GPU instances running overnight without a scheduled shutdown — expensive!
❌ Never deploy a model without a /health endpoint — you won't know when it crashes.
🗂️ The Complete End-to-End Architecture
┌────────────────────────────────────────────────────────────────┐
│ TRAINING PIPELINE │
│ │
│ HuggingFace Dataset │
│ │ │
│ ▼ │
│ Data Cleaning → Data Validation → Train/Val Split │
│ │ │
│ ▼ │
│ OCI GPU Cluster (H100 x 8) │
│ DistilBERT Fine-Tuning (3 epochs) │
│ │ │
│ ▼ │
│ MLflow Experiment Tracking │
│ │ │
│ ▼ │
│ Model Evaluation → F1: 0.87 ✅ │
│ │ │
│ ▼ │
│ Save Model → OCI Object Storage │
└───────────────────────┬────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ DEPLOYMENT PIPELINE │
│ │
│ Docker Image (serve.py + model + libraries) │
│ │ │
│ ▼ │
│ Oracle Container Registry (OCIR) │
│ │ │
│ ▼ │
│ OCI Kubernetes Engine (OKE) — 3 replicas │
│ │ │
│ ▼ │
│ OCI API Gateway → https://api.anomalydetector.com/detect │
│ │ │
│ ▼ │
│ OCI Monitoring + Alarms + MLflow Dashboards │
└────────────────────────────────────────────────────────────────┘
│
▼
User sends: "CRITICAL: Memory 99% — OOM kill"
Gets back: "ANOMALOUS ⚠️ — Confidence: 97.4%"
Happy fine-tuning! 🧠✨
Comments
Post a Comment