Skip to main content

AutoML in Machine Learning (OCI)

Calculating read time…

AutoML in Oracle Machine Learning (OML) works like a master chef for your data. Imagine you are cooking for a huge party, but you don't know which recipe will taste best, how much salt to add, or which ingredients actually matter. Now imagine a master chef who tastes ten different recipes for you, tweaks the spices automatically, and hands you back the single best dish — already plated. 🍽️

That master chef, for Machine Learning, is called AutoML. And inside Oracle Machine Learning (OML), this chef works right inside your database, never needing your data to travel anywhere else. 🧠🗄️

This is a long, detailed guide. Grab a coffee ☕ — by the end, you will understand not just what AutoML does, but why each of its four stages exists, when to trust it, and how to avoid the mistakes that trip up most beginners.

AutoML in Oracle Machine Learning: algorithm selection, feature selection, and model tuning inside the database

🧠 What Is AutoML, Really?

Building a good Machine Learning model by hand involves dozens of small decisions: which algorithm to use, which columns actually matter, what settings to use for that algorithm, and how to compare several candidate models fairly.

A skilled data scientist might spend days trying different combinations before landing on a model that performs well. Most of that time isn't spent on brilliant insight — it's spent on repetitive trial and error. 🔁

AutoML automates that repetitive trial and error. It systematically tries algorithms, narrows down which columns matter, tunes the settings, and ranks the results — so a beginner can get a strong model in minutes, and an expert can skip the boring parts and focus on judgment calls.

💡 Real-World Analogy:

Imagine you're allowed to hire ten chefs to each cook their own version of a dish, using slightly different spice combinations. You don't taste each dish yourself — instead, you have an expert food critic who tastes all ten, ranks them, and tells you exactly which one to serve, and why. AutoML is that food critic for your Machine Learning models. 👨‍🍳

Why This Actually Matters For Enterprises

In a big company, the bottleneck to using Machine Learning is rarely the algorithm itself — it's the number of skilled data scientists available to build and tune models for every single business problem. There are always more problems than there are experts to solve them.

AutoML doesn't replace data scientists. It multiplies their reach. A single data scientist can now supervise AutoML running across dozens of business problems at once, instead of manually tuning each one from scratch — freeing up their time for the problems that truly need human judgment, like fairness checks and business context. 👩‍💻


🧰 The Four Stages of AutoML in Oracle Machine Learning

Oracle's AutoML isn't one single black-box button. It's built from four distinct, inspectable stages, and you can use each one independently or let them run together. Understanding each stage separately is what turns you from a "user" into someone who truly understands what's happening under the hood.

Stage OML4Py Class What It Answers
1. Algorithm Selection oml.automl.AlgorithmSelection "Out of all available algorithms, which ones are likely to work best on my data?"
2. Adaptive Sampling (runs automatically inside the pipeline) "Do I really need to scan the whole huge table, or is a smart sample enough to decide?"
3. Feature Selection oml.automl.FeatureSelection "Which columns actually help predict the answer, and which are just noise?"
4. Model Tuning oml.automl.ModelTuning "For this one algorithm, what are the very best settings to use?"
5. Model Selection oml.automl.ModelSelection "Run all of the above together, and just hand me the single best final model."
✅ DO Remember:

You don't have to use all four stages every time. Sometimes you already know your algorithm, and you only want Feature Selection or Model Tuning. Other times, you just want the whole pipeline to run end-to-end using Model Selection. Oracle lets you pick the level of control you need. 🎛️

🗺️ Big Picture — How Does AutoML Flow End To End?

Diagram of the OML4Py AutoML pipeline: algorithm selection, feature selection, model tuning, ranked model list, and MLX explainability
  ┌───────────────┐   ┌──────────────────┐   ┌───────────────────┐   ┌──────────────────┐
  │  Your Table   │──►│ Algorithm        │──►│ Feature Selection │──►│  Model Tuning    │
  │ (raw data)    │   │ Selection        │   │ (drop noisy cols) │   │ (best settings)  │
  └───────────────┘   └──────────────────┘   └───────────────────┘   └────────┬─────────┘
                                                                                │
                                                                    ┌───────────▼──────────┐
                                                                    │   Ranked Model List   │
                                                                    │ → Best Model Chosen   │
                                                                    └───────────────────────┘

Simple flow: Raw data in → algorithms ranked → noisy columns dropped → best settings found → one winning model handed back to you. 🎯


🏗️ Step 1 — Set Up Your Environment

Step 1a: Get Access To Oracle Machine Learning

  • Create a free Autonomous Database (or Oracle AI Database Free) at cloud.oracle.com
  • Open the OML Notebooks interface, or install the OML4Py client locally
  • Both paths give you the exact same AutoML engine, running inside the database itself ✅

Step 1b: Connect Securely

📌 What does the code below do?

This code reads your database password from an environment variable instead of typing it directly into the script, then opens a connection with AutoML support switched on. 🔐
import os
import oml

db_password = os.environ["ORACLE_DB_PASSWORD"]

oml.connect(
    user="ml_user",
    password=db_password,
    dsn="mydb_high",
    automl=True
)

print("✅ Connected, AutoML engine ready.")
❌ DON'T do this:

Never hardcode your password directly in a notebook you plan to share or push to a code repository. 🔒

Step 1c: Prepare Your Training Data

📌 What does the code below do?

This code points a Python proxy object at a table already sitting inside your database, and then splits it into a training portion and a testing portion — like separating your study notes from the exam questions, so you can honestly check how well you learned. 📚
import oml

data = oml.sync(table="loan_applications")

# Split into 80% training data and 20% testing data
train, test = data.split(ratio=(0.8, 0.2), seed=42)

print(f"Training rows: {train.shape[0]}")
print(f"Testing rows : {test.shape[0]}")

Output:

Training rows: 8000
Testing rows : 2000

🥇 Step 2 — Algorithm Selection: "Which Recipe Might Win?"

Before training a single full model, it helps to know which algorithms are even worth trying. Algorithm Selection looks at the shape and characteristics of your data, and quickly ranks which algorithms are likely to perform best — without fully training all of them yet.

💡 Real-World Analogy:

Instead of asking ten chefs to cook a full three-course meal each, you first ask them to describe their recipe idea. An experienced critic can often guess which recipes are promising just from the description, before anyone even turns on the stove. That saves enormous time. 🍳
📌 What does the code below do?

This code asks AutoML to rank all supported classification algorithms for our loan default prediction problem, based on the accuracy score metric, and shows the top few candidates along with a predicted score for each — before fully training anything. 🏁
from oml import automl

algo_selector = automl.AlgorithmSelection(
    mining_function="classification",
    score_metric="accuracy"
)

ranked_algorithms = algo_selector.select(
    train, case_id="application_id", target="defaulted"
)

for algo, predicted_score in ranked_algorithms:
    print(f"{algo:20} predicted score: {predicted_score:.3f}")

Output:

SVM_GAUSSIAN         predicted score: 0.912
RANDOM_FOREST         predicted score: 0.905
NEURAL_NETWORK        predicted score: 0.898
GLM_CLASSIFICATION    predicted score: 0.861

SVM with a Gaussian kernel comes out on top — and now we know exactly where to focus our remaining time, instead of guessing. 🎯


✂️ Step 3 — Feature Selection: "Which Ingredients Actually Matter?"

Real enterprise tables often have 40, 60, even 100+ columns. Most of them don't actually help predict the outcome — they're just noise that slows training down and can even confuse the model. Feature Selection finds the small subset of columns that truly carry signal.

💡 Real-World Analogy:

Imagine baking a cake with 40 possible ingredients on your counter, but only 5 of them actually affect the taste — the rest are just sitting there taking up space. Feature Selection is the process of clearing the counter down to just those 5 ingredients that matter. 🎂
📌 What does the code below do?

This code uses the winning algorithm from Step 2 (SVM Gaussian) and asks Feature Selection to figure out which columns actually help predict loan default, then prints only that shortlist of useful columns. 📋
fs = automl.FeatureSelection(
    mining_function="classification",
    score_metric="accuracy"
)

selected_features = fs.reduce(
    train,
    case_id="application_id",
    target="defaulted",
    alg_name="svm_gaussian"
)

print("Columns AutoML kept:")
print(list(selected_features))

Output:

Columns AutoML kept:
['credit_score', 'monthly_income', 'existing_debt_ratio', 'employment_years', 'loan_amount']

Out of maybe 45 original columns, only 5 actually mattered — cutting training time significantly and often improving accuracy, since the model no longer gets distracted by irrelevant noise. 🧹

✅ DO Remember:

Fewer, more meaningful columns often produce a model that generalizes better to new data — this isn't just about speed, it genuinely helps model quality by reducing overfitting. 📐

🎛️ Step 4 — Model Tuning: "What's The Best Recipe Ratio?"

Even the right algorithm has "knobs" you can turn — called hyperparameters — like how many decision trees to grow, or how strict a boundary should be. Model Tuning automatically searches through these knobs to find the best combination.

💡 Real-World Analogy:

You've picked the recipe (algorithm) and the ingredients (features). Now you need to know: how much salt? How long to bake? Model Tuning tries different amounts automatically and keeps the combination that tasted best, instead of you guessing through trial and error. 🧂
📌 What does the code below do?

This code takes our chosen algorithm and the shortlisted features, then automatically searches for the best hyperparameter settings, such as the SVM's regularization strength. It returns a fully tuned model, ready to use. 🔧
mt = automl.ModelTuning(
    mining_function="classification",
    score_metric="accuracy"
)

tuning_results = mt.tune(
    "svm_gaussian",
    train[selected_features + ["defaulted"]],
    case_id="application_id",
    target="defaulted"
)

best_settings = tuning_results["best_model"]
print("Best hyperparameters found:")
print(tuning_results["all_evals"].head())

Output:

Best hyperparameters found:
   SVMS_COMPLEXITY_FACTOR   SVMS_KERNEL_FUNCTION   SCORE
0  1.35                     GAUSSIAN               0.931
1  0.80                     GAUSSIAN               0.918
2  2.10                     GAUSSIAN               0.904

Notice the accuracy jumped from 0.912 (Step 2's rough estimate) to 0.931 after real tuning — that gap is exactly what proper hyperparameter search buys you. 📈


🏆 Step 5 — Model Selection: Running The Whole Pipeline At Once

If you don't want to run each stage manually, Model Selection runs Algorithm Selection, Feature Selection, and Model Tuning together, and simply hands you back the single best model at the end.

📌 What does the code below do?

This code runs the entire AutoML pipeline in one call — trying multiple algorithms, narrowing down features, and tuning settings — then returns the winning, ready-to-use model object, plus its final score. 🎯
model_selection = automl.ModelSelection(
    mining_function="classification",
    score_metric="accuracy"
)

best_model = model_selection.select(
    train, case_id="application_id", target="defaulted"
)

print("Winning algorithm:", best_model.algorithm_name)
print("Final tuned score :", best_model.score)

Output:

Winning algorithm: SVM_GAUSSIAN
Final tuned score : 0.931

Notice this matches exactly what we found manually across Steps 2-4 — because that's literally what Model Selection is doing behind the scenes, just automatically and much faster. ⚡

Step 5a: Evaluate The Winning Model Honestly

📌 What does the code below do?

This code checks the winning model's performance on the test data it has never seen before, which is the honest, real-world measure of whether the model is actually good. 🧪
predictions = best_model.predict(test)
accuracy_on_unseen_data = (predictions["PREDICTION"] == test["defaulted"]).mean()

print(f"Accuracy on unseen test data: {accuracy_on_unseen_data:.3f}")

Output:

Accuracy on unseen test data: 0.924
✅ DO Remember:

A small drop from 0.931 (training-time estimate) to 0.924 (real test accuracy) is completely normal and healthy. If you ever see a huge drop, that's a warning sign of overfitting — the model memorized the training data instead of learning general patterns. 🚩

🖱️ Step 6 — The No-Code Path: AutoML UI

Not everyone on your team writes Python. Oracle Machine Learning also ships a no-code AutoML UI, built directly into the OML Notebooks interface, for analysts and business users who think in spreadsheets, not scripts.

How the No-Code Flow Works

  1. Step A: Open AutoML UI and click "Create Experiment"
  2. Step B: Choose your data source table and the column you want to predict
  3. Step C: Click "Run" — AutoML performs algorithm selection, feature selection, and tuning automatically
  4. Step D: Review a ranked leaderboard of models, sorted by your chosen quality metric
  5. Step E: Click "Generate Notebook" on your favorite model to get the exact OML4Py code that produced it
✅ Why This Matters:

This "Generate Notebook" step is genuinely clever — a business analyst can explore models visually, then hand the auto-generated Python code straight to an engineering team for productionizing, with zero manual re-implementation needed. 🤝

🔍 Step 7 — Understanding Why The Model Decided What It Decided (MLX)

A model that predicts well but can't explain itself is risky in regulated industries. Oracle's Machine Learning Explainability (MLX) tools help you see why a model made a particular prediction, not just what it predicted.

📌 What does the code below do?

This code asks the model to explain one specific prediction — showing which factors pushed this particular applicant toward a "default" prediction, and by how much. This turns a black-box answer into something a loan officer can actually justify to a customer. ⚖️
from oml import mlx

explainer = mlx.LocalExplainer(model=best_model, train_data=train)

explanation = explainer.explain(
    row=test.iloc[0],
    target="defaulted"
)

print(explanation.top_contributing_features())

Output:

Feature                  Contribution
------------------------ ------------
existing_debt_ratio       +0.42 (increases default risk)
credit_score              -0.31 (decreases default risk)
employment_years          -0.10 (decreases default risk)

Now instead of just "the model says 78% chance of default," you can tell a customer exactly which factors mattered most — building trust and satisfying regulatory explainability requirements. 📋


🏢 A Named Enterprise Scenario: "GreenLeaf Insurance"

Let's make this concrete. GreenLeaf Insurance (a fictional example) needed to predict which policy applications were likely to result in early claims, across 12 different insurance product lines. Before AutoML, their small data science team could only properly support 3 of those product lines with hand-tuned models — the rest used a single generic model that performed poorly everywhere.

After adopting AutoML through Oracle Machine Learning, their new workflow looked like this:

  • Each of the 12 product lines now gets its own AutoML-generated model, tuned specifically to its data
  • Data scientists spend their time reviewing MLX explanations for fairness and compliance, not tuning hyperparameters by hand
  • New product lines can get a working baseline model within a single afternoon, instead of a multi-week project
  • Overall claim-prediction accuracy across all product lines improved by roughly 18%, largely because every line finally got a model actually tuned for its own data
✅ The Real Lesson:

The improvement wasn't from one smarter algorithm — it came from finally being able to give every single problem its own properly-tuned model, because AutoML removed the bottleneck of human time. This is the true enterprise value of AutoML: scale, not just accuracy. 📈

⚖️ Manual Tuning vs AutoML: An Honest Comparison

Aspect Manual Tuning By An Expert AutoML
Time To First Model Days to weeks Minutes to hours
Consistency Varies by individual skill and available time Consistent, repeatable process every time
Best For Novel problems needing deep domain judgment Standard classification/regression at scale across many problems
Human Role Does the tuning directly Reviews results, checks fairness, makes the final judgment call

🏆 Best Practices

  • 🎯 Always keep a held-out test set — AutoML's own internal scores are estimates, not proof
  • ⚖️ Choose your score metric carefully — accuracy is misleading on imbalanced data; consider F1 or recall
  • ✂️ Run Feature Selection even if you're confident in your columns — it often finds surprises
  • 🔍 Always pair AutoML with MLX explainability in regulated industries, before deploying to production
  • 🔁 Re-run AutoML periodically — the best algorithm for your data today may not be the best in a year
  • 📝 Use "Generate Notebook" in AutoML UI to hand business-friendly experiments to engineering teams cleanly
❌ Common Mistakes to Avoid:

  • Do NOT trust AutoML's training-time score as the final answer — always validate on unseen test data
  • Do NOT skip explainability just because the model is convenient — regulators and customers may ask "why"
  • Do NOT assume AutoML understands your business context — it optimizes the metric you give it, nothing more
  • Do NOT let AutoML pick a technically "better" but non-compliant algorithm for regulated decisions without a compliance review

🌍 Real-World Use Cases

  • 🏦 Banking: Credit risk scoring models, tuned separately per loan product
  • 🏥 Healthcare: Readmission risk models, retrained automatically as patient patterns shift
  • 🛒 Retail: Demand forecasting models generated per product category, at scale
  • 📞 Telecom: Churn prediction tuned separately for prepaid vs postpaid customers
  • 🏭 Manufacturing: Equipment failure prediction models, one per machine type
  • ⚖️ Insurance: Claims risk models with built-in explainability for regulatory review

  • AutoML Meets Vector Data — AutoML pipelines increasingly consume embeddings from AI Vector Search as engineered features, blending semantic meaning with classic tabular ML
  • Explainability As A Default, Not An Add-On — MLX-style explanations are increasingly generated automatically alongside every AutoML model, not requested separately
  • No-Code Democratization — AutoML UI continues to close the gap between business analysts and engineering teams through the "Generate Notebook" hand-off
  • AutoML Across More Problem Types — the same automated approach now extends beyond classification and regression into time series and anomaly detection tuning

📝 Quick Summary — What We Learned

  • What AutoML is → An automated system that selects algorithms, features, and settings so you don't have to do it all by hand
  • Algorithm Selection → Quickly ranks which algorithms are worth fully training
  • Feature Selection → Narrows down to only the columns that truly help predict the target
  • Model Tuning → Automatically searches for the best hyperparameter settings
  • Model Selection → Runs the whole pipeline together and hands back the single best model
  • AutoML UI → A no-code path for business analysts, with a "Generate Notebook" bridge to engineering
  • MLX Explainability → Explains why a model made a specific prediction, critical for regulated industries
  • Enterprise Pattern → Use AutoML to scale properly-tuned models across many problems, and keep human judgment focused on fairness, compliance, and business context

❓ Troubleshooting & Frequently Asked Questions

Does AutoML always find the globally best possible model?

No — it finds a strong model efficiently by intelligently searching a large space of options, but it doesn't exhaustively try every possible combination. For most enterprise problems, this efficient search is more than good enough, and dramatically faster than exhaustive search.

Can I restrict AutoML to only try certain algorithms?

Yes. You can pass a specific list of algorithms to consider, which is useful when you already know certain algorithms are required or forbidden for compliance reasons in your industry.

Is AutoML slower than just picking one algorithm manually?

For a single quick experiment, yes, manual can be faster. But across dozens of business problems, AutoML's consistency and speed per-problem wins decisively, especially when data scientist time is the true bottleneck, as it is at most enterprises.

Does using AutoML mean I don't need a data scientist anymore?

No. AutoML removes repetitive tuning work, but judgment calls — fairness checks, business context, deciding which metric actually matters, and validating explainability — still need a human expert.

Happy building! 🧠✨

Comments