OCI Forecasting — the time-series forecasting capability built into Oracle Cloud Infrastructure — works like a weather forecaster who doesn't just look at today's temperature, but studies years of past weather, notices that rain always follows a pressure drop, and tells you tomorrow's forecast along with exactly how confident they are. 🌦️
Today, that capability ships as the Forecasting Operator, part of Oracle's Accelerated Data Science (ADS) library inside OCI Data Science — a low-code, YAML-driven tool that competes several forecasting frameworks against each other on your actual data and keeps the winner, without you needing a data scientist on staff to hand-build a model. 📈🧠
This is a long, detailed, platinum-grade guide. Grab a coffee ☕ — by the end, you will understand not just what this capability does, but exactly how to configure it, how it explains its own predictions, and the mistakes that trip up most beginners.
- What Is OCI Forecasting?
- What Makes It Different?
- Single-Series vs. Regressor-Enriched Forecasting
- Big Picture — How Does It All Flow?
- Step 1 — Set Up Your Environment
- Step 2 — Initialize the Forecasting Operator
- Step 3 — Your First Forecast Spec
- Step 4 — Run It and Read the Results
- Step 5 — Model Selection, Per Series
- Step 6 — Adding External Regressors
- Step 7 — Explainability
- Step 8 — What-If Analysis & Deployment
- Step 9 — A Complete End-to-End Example
- Enterprise Architecture
- A Named Enterprise Scenario
- Best Practices & Common Mistakes
- Real-World Use Cases
- What's New in 2026
- Quick Summary
- FAQ
🧠 What Is OCI Forecasting?
Most businesses forecast the future in one of two painful ways: someone in Excel manually drags a trend line across a chart and guesses, or a data science team spends weeks hand-building a custom statistical model from scratch for every single metric that needs predicting.
The ADS Forecasting Operator skips both of those problems. You describe your data and your goal in a short YAML file, and the operator trains and compares several forecasting frameworks — statistical, machine learning, and deep learning — automatically selecting whichever one fits your actual historical pattern best.
Think of it like this 👇
- You have 2 years of daily sales history sitting in a CSV 📊
- You describe it in a YAML file and say: "Forecast the next 30 days"
- The operator studies patterns, seasonality, and trends automatically
- It hands back a day-by-day forecast, plus a confidence range for each day ✅
Imagine hiring several fortune tellers — a statistician, a classic time-series expert, and a couple of deep-learning specialists — and asking all of them to predict the same number. Instead of picking one at random, an impartial judge checks which of them was most accurate on past data, and only tells you the answer from the best performer. That judge is the Forecasting Operator's automatic model selection. 🔮
Why "No Data Science Expertise Needed" Is a Big Deal
Building a good forecasting model usually means knowing which algorithm fits which kind of pattern — should you use exponential smoothing, ARIMA, Prophet, or a neural network like NeuralProphet? Getting this wrong produces confidently wrong forecasts, which are often worse than no forecast at all.
Set the operator's model
field to auto
and it evaluates candidate frameworks against your data and keeps the one
with the strongest backtested accuracy — in plain words, it tries several
approaches on your actual numbers and quietly picks the winner for you.
🌟 What Makes OCI Forecasting Different?
| Feature | What It Means For You |
|---|---|
| Automated Model Selection | Setting model: auto lets the operator pick the best-fitting framework for you, instead of you guessing between Prophet, ARIMA, NeuralProphet, AutoMLX, or AutoTS |
| Confidence Intervals | Every forecast row ships with upper and lower bounds, sized to the confidence width you configure |
| Explainability | Setting generate_explanations: true produces both global and local explanation files, showing which factors drove which predictions |
| External Regressors | An optional additional_data source lets promotions, prices, or weather feed directly into the forecast |
| Per-Series Auto Selection | When forecasting many series at once (e.g. one per store), the operator can recommend a different best-fit model for each series independently |
| What-If Analysis | Trained models can be packaged and optionally deployed as a live endpoint, so planners can test scenarios in real time |
🔀 Single-Series vs. Regressor-Enriched Forecasting
Before writing any YAML, it helps to understand the two shapes a forecast can take, since choosing the right one determines what data you need to prepare.
| Mode | What It Uses | Best For |
|---|---|---|
| Single-series (historical_data only) | Only the target metric's own past history (e.g. past sales predicting future sales) | Simple metrics with strong internal patterns, like website traffic or basic seasonal sales |
| Regressor-enriched (+ additional_data) | The target metric plus known external drivers such as promotions, prices, or weather | Retail demand, revenue affected by marketing campaigns, energy usage affected by temperature |
If you know a specific factor genuinely drives your metric — like a promotion boosting sales — feeding it in through
additional_data almost
always improves accuracy over a plain single-series forecast. The catch: you
must actually know (or credibly plan) that driver's value for every day in
the forecast horizon, not just historically. 📐
🗺️ Big Picture — How Does It All Flow?
┌────────────────────┐ ┌───────────────────────┐ ┌─────────────────────────┐
│ historical_data │───►│ forecast.yaml spec: │───►│ ads operator run │
│ (+ additional_data)│ │ frameworks compete │ │ Best model wins │
└────────────────────┘ └───────────────────────┘ └───────────┬─────────────┘
│
┌────────────────────────▼───────────────────────┐
│ forecast.csv + explanations + report.html │
└────────────────────────┬───────────────────────┘
│
┌────────────▼─────────────┐
│ Dashboard / Planning App │
└───────────────────────────┘
Simple flow: Describe your data in YAML → frameworks compete → best one wins → you get a forecast, a confidence range, and a "why" behind it. 🎯
🏗️ Step 1 — Set Up Your OCI Environment
Step 1a: Provision a Place to Run It
- Go to cloud.oracle.com and sign in (or create a free account)
- Navigate to Analytics & AI → Data Science and create a project and a notebook session
- Ask your administrator to grant your group access with a policy ✅
It tells OCI: "Members of MyDevelopers group are allowed to manage Data Science resources — notebooks, projects, models — in this tenancy." Without this, provisioning a notebook session will fail with a permission error. 📋
allow group MyDevelopers to manage data-science-family in tenancy
Step 1b: Install the Forecasting Operator
This installs the Accelerated Data Science library along with its forecasting extras (Prophet, ARIMA, NeuralProphet, and friends). If you're working inside an OCI Data Science notebook session, this is often already available through the provided environment — run this only if you're setting up locally or a custom environment. 🔧
python3 -m pip install "oracle-ads[forecast]"
Never share your
~/.oci/config file or private key with anyone.
Keep it as secret as a house key. 🔒
📁 Step 2 — Initialize the Forecasting Operator
Instead of a "project + data asset" API call, the operator starts from a scaffolded folder generated by a single CLI command.
This creates a new folder containing a starter
forecast.yaml file
— a bare-bones template with just enough fields to run a first forecast. You
edit this file to point at your real data, then run it. Nothing is
forecasted yet at this step. 📂
ads operator init --type forecast --output daily-sales-forecast
What you'll find inside the new folder:
daily-sales-forecast/ └── forecast.yaml # starter spec — edit this next
📊 Step 3 — Your First Forecast Spec
This tells the operator: "Here's two years of daily revenue, in a file where the date column is called `sales_date` and the value column is called `daily_revenue` — predict the next 30 days, and figure out which modeling framework fits best on your own." Every field beyond
datetime_column, historical_data,
horizon, and target_column is optional, but adding
a few makes the output far more useful. 🎯
kind: operator
type: forecast
version: v1
spec:
datetime_column:
name: sales_date
format: "%Y-%m-%d"
historical_data:
url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
target_column: daily_revenue
horizon: 30
model: auto
metric: smape
generate_explanations: true
output_directory:
url: results
Save this as forecast.yaml inside the folder from Step 2.
📥 Step 4 — Run It and Read the Results
This is the one command that actually trains everything: it loads your historical data, tries the candidate frameworks, keeps the best performer, generates a forecast, and writes every output file into the
output_directory you configured (or a folder called
results if you didn't set one). ⏳
ads operator run -f forecast.yaml
What lands in your output folder when it finishes:
forecast.csv— date, forecast value, lower bound, upper bound, per rowmetrics.csv— the backtested error score for the winning modelglobal_explanations.csv/local_explanations.csv— covered in Step 7report.html— a shareable visual summary a business planner can open directly
This reads
forecast.csv with pandas and prints the first 5
forecasted days next to their confidence range — the exact same file your
dashboard or planning app would consume. 📅
import pandas as pd
forecast_df = pd.read_csv("results/forecast.csv")
print(forecast_df[["Date", "forecast_value", "lower_bound", "upper_bound"]].head())
Output:
Date forecast_value lower_bound upper_bound 2026-07-06 48230.0 44100.0 52360.0 2026-07-07 51940.0 47500.0 56380.0 2026-07-08 49870.0 45600.0 54140.0 2026-07-09 53210.0 48800.0 57620.0 2026-07-10 55480.0 50900.0 60060.0
Notice every prediction comes with a range, not just one number — that range tells your business exactly how much uncertainty to plan around. 📊
🏆 Step 5 — Model Selection, Per Series
The model field accepts a specific framework name
(prophet, arima, neuralprophet,
automlx, autots) if you already know what fits your
data, or auto to let the operator decide. When you're
forecasting many related series at once — say, one per store —
auto-selection can go a level deeper and recommend a different best-fit
framework independently for each individual series, since a steady seller and
a spiky, promotion-driven product rarely share the same ideal model.
Adding
target_category_columns tells the operator your CSV
actually contains many interleaved series (one per store), not just one —
it will train and forecast each store independently, in the same run. 🧩
spec:
historical_data:
url: oci://sales-data-bucket@your_namespace/all_stores_sales.csv
target_category_columns:
- store_id
target_column: daily_revenue
datetime_column:
name: sales_date
horizon: 30
model: auto
Reading the winning framework per series from metrics.csv:
import pandas as pd
metrics_df = pd.read_csv("results/metrics.csv")
print(metrics_df[["store_id", "selected_model", "smape"]].head())
store_id selected_model smape 101 neuralprophet 4.12 102 arima 6.03 103 prophet 5.44
A deep learning model won for store 101 — but a classic statistical method won for store 102. This is exactly why per-series automatic selection matters at real retail scale: no single algorithm is right for every store. 🥇
🔗 Step 6 — Adding External Regressors
Sales rarely move in isolation — promotions, prices, and even weather can be
genuine drivers. The operator's additional_data source lets
those known, plannable factors feed directly into the forecast.
A single-series forecast is like predicting tomorrow's temperature only from yesterday's temperature. Adding regressors is like also checking whether a storm system is approaching — it uses more relevant information to make a smarter guess. 🌩️
This trains the same forecast as before, but this time hands the operator a second file with an
on_promotion and average_price
column for every date in both the historical range and the future
horizon. The model learns how these factors historically influenced sales,
and factors that relationship into its future prediction. Critically, this
second file must already contain values reaching across the whole forecast
horizon — you have to know or credibly plan these values in advance. 🧩
spec:
datetime_column:
name: sales_date
format: "%Y-%m-%d"
historical_data:
url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
additional_data:
url: oci://sales-data-bucket@your_namespace/promo_and_price_plan.csv
target_column: daily_revenue
horizon: 30
model: auto
generate_explanations: true
output_directory:
url: results
ads operator run -f forecast.yaml
🔍 Step 7 — Understanding Why the Forecast Looks the Way It Does
Setting generate_explanations: true produces two extra files
that turn "trust me" into "here's exactly why."
This reads
global_explanations.csv — the file that describes
which factors influenced the whole forecast the most overall — so a business
planner can see why the model expects a spike or dip, not just
that it expects one. There's a matching
local_explanations.csv that breaks this down per individual
prediction, not just in aggregate. 📋
import pandas as pd
global_df = pd.read_csv("results/global_explanations.csv")
print(global_df.sort_values("importance_score", ascending=False).head())
feature importance_score on_promotion 0.47 average_price 0.31 day_of_week 0.22
Promotions turn out to be the single biggest driver of revenue swings — more influential than even price changes. That's a genuinely useful business insight, not just a number. 💡
🔁 Step 8 — What-If Analysis & Model Deployment
A forecast trained once and never revisited slowly goes stale as real business conditions shift. The operator's newer What-If Analysis capability packages a trained model into the OCI Model Catalog and, optionally, deploys it as a live endpoint — so planners can test "what if we ran a bigger promotion next month?" scenarios without retraining from scratch.
Adding a
what_if_analysis block tells the operator to save the
trained model bundle to the Model Catalog after this run. Including a
model_deployment block goes one step further and has ADS
automatically stand up a live HTTP endpoint, complete with logging and
autoscaling — so a planning app can POST a hypothetical
additional_data scenario and get a new forecast back instantly. 🗓️
spec:
# ...same fields as Step 6...
what_if_analysis:
model_deployment:
display_name: "sales-forecast-whatif-endpoint"
auto_scaling:
min_instances: 1
max_instances: 3
What you'll find after the run:
results/deployment_info.json → model catalog OCID → model deployment OCID → live endpoint URL
From here, updating the promotion or price assumptions in a fresh
additional_data file and POSTing it to that endpoint returns an
updated forecast in real time — without a full retrain. 🚀
🧵 Step 9 — The Complete End-to-End Example, Start to Finish
Let's stitch everything together into one real, runnable spec plus the two commands that take it from an empty folder to a printed forecast table. 🧵
- Scaffolds a new forecast project folder
- Points at your historical sales data already sitting in Object Storage
- Requests a 30-day forecast with automatic model selection and explanations
- Runs the operator, which trains, compares, and picks the best model
- Reads back the first 5 days of the forecast, with confidence ranges
# ──── 1. Scaffold the project ───────────────────────────────────────────────
$ ads operator init --type forecast --output daily-sales-forecast
$ cd daily-sales-forecast
# ──── 2. Edit forecast.yaml ─────────────────────────────────────────────────
cat > forecast.yaml <<'EOF'
kind: operator
type: forecast
version: v1
spec:
datetime_column:
name: sales_date
format: "%Y-%m-%d"
historical_data:
url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
target_column: daily_revenue
horizon: 30
model: auto
metric: smape
generate_explanations: true
output_directory:
url: results
EOF
# ──── 3. Run the forecast ───────────────────────────────────────────────────
$ ads operator run -f forecast.yaml
# ──── 4. Read the results ───────────────────────────────────────────────────
import pandas as pd
df = pd.read_csv("results/forecast.csv")
print(df[["Date", "forecast_value", "lower_bound", "upper_bound"]].head())
Output:
Date forecast_value lower_bound upper_bound 2026-07-06 48230.0 44100.0 52360.0 2026-07-07 51940.0 47500.0 56380.0 2026-07-08 49870.0 45600.0 54140.0 2026-07-09 53210.0 48800.0 57620.0 2026-07-10 55480.0 50900.0 60060.0
A complete, working pipeline that goes from raw historical sales data, to a YAML-configured forecast run, to a 30-day forecast with confidence ranges — in one short spec plus two CLI commands, using the real ADS Forecasting Operator. 🏁
🏛️ Enterprise Architecture — Where This Fits
- 🗄️ Object Storage / Data Lake — holds historical time series data
- 🔄 OCI Data Integration — moves and refreshes historical data on a schedule
- 📈 OCI Data Science + ADS Forecasting Operator — trains, selects the best model per series, and produces forecasts
- 📦 Model Catalog & Model Deployment — stores trained models and optionally serves them for live What-If scenarios
- 📊 Oracle Analytics Cloud — visualizes forecasts and confidence ranges for business users
- 🗂️ Business Applications — supply chain, inventory, or workforce planning tools consuming the forecast output
[ Historical Data In Object Storage ] -> [ ADS Forecasting Operator: Train + Select Best Model ]
-> [ forecast.csv + Confidence Interval + Explanations ]
-> [ Model Catalog (+ optional live Deployment) ]
-> [ Oracle Analytics Cloud Dashboard ]
-> [ Supply Chain / Planning App Consumes Forecast Output ]
🏢 A Named Enterprise Scenario: "BrightMart Retail"
Let's make this concrete. BrightMart Retail (a fictional example) runs 120 stores and previously forecasted demand using a single company-wide spreadsheet model, updated manually once a month by a small planning team.
- Stockouts during promotional weekends were common, since the old model didn't account for promotions
- Overstock of slow-moving items tied up warehouse space and cash
- The planning team could only realistically update forecasts for the 20 highest-revenue stores
After adopting the ADS Forecasting Operator with a per-store, promotion-aware spec:
- All 120 stores now get their own forecast in a single run, using
target_category_columnsto keep every store's series independent - Promotional periods are factored in directly, since planned promotion and price data feeds in through
additional_data - Stockouts during promotions dropped significantly, since inventory could be pre-positioned ahead of expected demand spikes
- The planning team now spends its time reviewing
global_explanations.csvand exceptions, not manually rebuilding spreadsheets - A live What-If deployment lets regional managers test "what if we extend the promotion a week?" without waiting for a full retrain
The win wasn't a smarter single model — it was being able to give every single store its own properly-tuned, automatically-refreshed forecast, something that was simply impossible for a small human team to do manually at that scale. 📈
🏆 Best Practices
- 📅 Feed in as much clean history as you have — more history generally helps the model learn seasonality correctly
- 🔗 Add known external factors via additional_data — promotions, price changes, and holidays often meaningfully improve accuracy
- 🎯 Always check the confidence interval, not just the single forecast number, when making planning decisions
- 🥇 Review metrics.csv per series — a close second-place model can be a useful sanity check, especially with
model: auto - 🔍 Use global_explanations.csv and local_explanations.csv to validate that the model's reasoning actually matches real business knowledge
- 🔁 Re-run on a schedule — stale forecasts quietly become inaccurate as real conditions shift
- Do NOT feed in very short histories (a few weeks) and expect the model to learn seasonality
- Do NOT ignore the confidence interval — planning around only the midpoint forecast hides real risk
- Do NOT put a factor in
additional_dataif you won't actually know its future value in advance — the file must cover the full horizon, not just history - Do NOT disable default preprocessing steps (missing-value imputation, outlier treatment) without understanding why — it can cause a framework to fail outright
- Do NOT forget to re-check forecasts after a major business disruption — the historical pattern the model learned may no longer apply
🌍 Real-World Use Cases
- 🛒 Retail: Product demand forecasting per store, accounting for promotions and seasonality
- 🏭 Manufacturing: Raw material and component demand planning
- 💰 Finance: Revenue and cash flow forecasting for budgeting cycles
- ☎️ Customer Support: Call center volume forecasting for staffing decisions
- ⚡ Utilities & Energy: Demand forecasting influenced by weather conditions
- 🖥️ IT Operations: Forecasting compute or storage capacity needs ahead of demand spikes
🚀 What's New and Trending in 2026
- Per-Series Model Selection at Scale — modern forecasting workloads increasingly select the best algorithm independently for each individual series, since the best-performing framework genuinely varies by frequency, horizon, and signal pattern
- Real-Time What-If Analysis — trained models packaged into the Model Catalog and optionally deployed live, so planners can test scenarios without a full retrain
- Forecasting Feeding Generative AI Narratives — explanation outputs increasingly feed directly into Generative AI, which drafts plain-language planning summaries for business stakeholders
- Explainability As Standard Practice — global and local influencer breakdowns are increasingly treated as a required output, not an optional extra, especially in regulated planning processes
📝 Quick Summary — What We Learned
- What OCI Forecasting is → The ADS Forecasting Operator, a low-code, YAML-driven capability inside OCI Data Science that delivers time-series forecasts via statistical, ML, and deep learning frameworks, with no data science expertise required
- Automated Model Selection → Set
model: autoto try multiple frameworks and keep the one with the strongest backtested accuracy, per series if needed - Single-Series vs. Regressor-Enriched → Use only the metric's own history, or enrich it with
additional_datalike promotions and price - Confidence Intervals → Every forecast comes with a range, not just a single guessed number
- Explainability →
global_explanations.csvandlocal_explanations.csvshow exactly why the model predicted what it did - What-If Analysis → Package and optionally deploy a model for live scenario testing, not just static reports
- Enterprise Pattern → Scale properly-tuned, explainable forecasts across every product, store, or metric that matters — not just the few a small team could handle manually
❓ Troubleshooting & Frequently Asked Questions
There's no hard minimum, but generally at least 2-3 full seasonal cycles (often 2 years for daily/weekly business data) gives the model enough signal to properly learn recurring patterns.
Yes — add the relevant column(s) to target_category_columns so the operator trains and forecasts a series per group, such as one per product or one per store, within the same run.
The most common cause is a mismatch between datetime_column's name/format and what's actually in your file, or an additional_data file that doesn't extend across the full forecast horizon. Check those two first before anything else.
It removes the need to hand-build classical forecasting models for routine business metrics. Highly specialized, novel forecasting problems may still benefit from custom modeling directly in OCI Data Science, but most enterprise demand and revenue forecasting fits comfortably within this operator.
Conceptually yes — same goal, same core ideas (auto model selection, confidence intervals, explainability). The 2021 announcement described a limited-availability standalone API; the generally available path today is the ADS Forecasting Operator inside OCI Data Science, which is what this guide covers.
Happy building! 📈✨
Comments
Post a Comment