Skip to main content

OCI Forecasting : Delivering time-series forecasts

Calculating read time…

OCI Forecasting — the time-series forecasting capability built into Oracle Cloud Infrastructure — works like a weather forecaster who doesn't just look at today's temperature, but studies years of past weather, notices that rain always follows a pressure drop, and tells you tomorrow's forecast along with exactly how confident they are. 🌦️

Today, that capability ships as the Forecasting Operator, part of Oracle's Accelerated Data Science (ADS) library inside OCI Data Science — a low-code, YAML-driven tool that competes several forecasting frameworks against each other on your actual data and keeps the winner, without you needing a data scientist on staff to hand-build a model. 📈🧠

This is a long, detailed, platinum-grade guide. Grab a coffee ☕ — by the end, you will understand not just what this capability does, but exactly how to configure it, how it explains its own predictions, and the mistakes that trip up most beginners.

OCI Forecasting with the ADS Forecasting Operator: automated model selection, confidence intervals, and explainability
⚠️ A Quick, Honest Note on Naming: Oracle originally previewed a standalone "OCI Forecasting" AI service back in 2021, available only through a limited-availability beta program. The generally available, production path today is the ADS Forecasting Operator, which lives inside OCI Data Science. Everything in this guide targets that real, currently-supported path — not the earlier beta API, which most readers can no longer access anyway.

🧠 What Is OCI Forecasting?

Most businesses forecast the future in one of two painful ways: someone in Excel manually drags a trend line across a chart and guesses, or a data science team spends weeks hand-building a custom statistical model from scratch for every single metric that needs predicting.

The ADS Forecasting Operator skips both of those problems. You describe your data and your goal in a short YAML file, and the operator trains and compares several forecasting frameworks — statistical, machine learning, and deep learning — automatically selecting whichever one fits your actual historical pattern best.

Think of it like this 👇

  • You have 2 years of daily sales history sitting in a CSV 📊
  • You describe it in a YAML file and say: "Forecast the next 30 days"
  • The operator studies patterns, seasonality, and trends automatically
  • It hands back a day-by-day forecast, plus a confidence range for each day ✅
💡 Real-World Analogy:

Imagine hiring several fortune tellers — a statistician, a classic time-series expert, and a couple of deep-learning specialists — and asking all of them to predict the same number. Instead of picking one at random, an impartial judge checks which of them was most accurate on past data, and only tells you the answer from the best performer. That judge is the Forecasting Operator's automatic model selection. 🔮

Why "No Data Science Expertise Needed" Is a Big Deal

Building a good forecasting model usually means knowing which algorithm fits which kind of pattern — should you use exponential smoothing, ARIMA, Prophet, or a neural network like NeuralProphet? Getting this wrong produces confidently wrong forecasts, which are often worse than no forecast at all.

Set the operator's model field to auto and it evaluates candidate frameworks against your data and keeps the one with the strongest backtested accuracy — in plain words, it tries several approaches on your actual numbers and quietly picks the winner for you.


🌟 What Makes OCI Forecasting Different?

Feature What It Means For You
Automated Model Selection Setting model: auto lets the operator pick the best-fitting framework for you, instead of you guessing between Prophet, ARIMA, NeuralProphet, AutoMLX, or AutoTS
Confidence Intervals Every forecast row ships with upper and lower bounds, sized to the confidence width you configure
Explainability Setting generate_explanations: true produces both global and local explanation files, showing which factors drove which predictions
External Regressors An optional additional_data source lets promotions, prices, or weather feed directly into the forecast
Per-Series Auto Selection When forecasting many series at once (e.g. one per store), the operator can recommend a different best-fit model for each series independently
What-If Analysis Trained models can be packaged and optionally deployed as a live endpoint, so planners can test scenarios in real time

🔀 Single-Series vs. Regressor-Enriched Forecasting

Before writing any YAML, it helps to understand the two shapes a forecast can take, since choosing the right one determines what data you need to prepare.

Mode What It Uses Best For
Single-series (historical_data only) Only the target metric's own past history (e.g. past sales predicting future sales) Simple metrics with strong internal patterns, like website traffic or basic seasonal sales
Regressor-enriched (+ additional_data) The target metric plus known external drivers such as promotions, prices, or weather Retail demand, revenue affected by marketing campaigns, energy usage affected by temperature
✅ DO Remember:

If you know a specific factor genuinely drives your metric — like a promotion boosting sales — feeding it in through additional_data almost always improves accuracy over a plain single-series forecast. The catch: you must actually know (or credibly plan) that driver's value for every day in the forecast horizon, not just historically. 📐

🗺️ Big Picture — How Does It All Flow?

Diagram of the ADS Forecasting Operator flow: historical and additional data, forecast.yaml spec, ads operator run, forecast and explanation outputs, model catalog, and dashboard
  ┌────────────────────┐    ┌───────────────────────┐    ┌─────────────────────────┐
  │  historical_data   │───►│  forecast.yaml spec:  │───►│  ads operator run       │
  │  (+ additional_data)│    │  frameworks compete   │    │  Best model wins        │
  └────────────────────┘    └───────────────────────┘    └───────────┬─────────────┘
                                                                       │
                                              ┌────────────────────────▼───────────────────────┐
                                              │  forecast.csv + explanations + report.html      │
                                              └────────────────────────┬───────────────────────┘
                                                                       │
                                                          ┌────────────▼─────────────┐
                                                          │  Dashboard / Planning App │
                                                          └───────────────────────────┘

Simple flow: Describe your data in YAML → frameworks compete → best one wins → you get a forecast, a confidence range, and a "why" behind it. 🎯


🏗️ Step 1 — Set Up Your OCI Environment

Step 1a: Provision a Place to Run It

  • Go to cloud.oracle.com and sign in (or create a free account)
  • Navigate to Analytics & AI → Data Science and create a project and a notebook session
  • Ask your administrator to grant your group access with a policy ✅
📌 What does the policy below do?

It tells OCI: "Members of MyDevelopers group are allowed to manage Data Science resources — notebooks, projects, models — in this tenancy." Without this, provisioning a notebook session will fail with a permission error. 📋
allow group MyDevelopers to manage data-science-family in tenancy

Step 1b: Install the Forecasting Operator

📌 What does the command below do?

This installs the Accelerated Data Science library along with its forecasting extras (Prophet, ARIMA, NeuralProphet, and friends). If you're working inside an OCI Data Science notebook session, this is often already available through the provided environment — run this only if you're setting up locally or a custom environment. 🔧
python3 -m pip install "oracle-ads[forecast]"
❌ DON'T do this:

Never share your ~/.oci/config file or private key with anyone. Keep it as secret as a house key. 🔒

📁 Step 2 — Initialize the Forecasting Operator

Instead of a "project + data asset" API call, the operator starts from a scaffolded folder generated by a single CLI command.

📌 What does the command below do?

This creates a new folder containing a starter forecast.yaml file — a bare-bones template with just enough fields to run a first forecast. You edit this file to point at your real data, then run it. Nothing is forecasted yet at this step. 📂
ads operator init --type forecast --output daily-sales-forecast

What you'll find inside the new folder:

daily-sales-forecast/
└── forecast.yaml   # starter spec — edit this next

📊 Step 3 — Your First Forecast Spec

📌 What does the YAML below do?

This tells the operator: "Here's two years of daily revenue, in a file where the date column is called `sales_date` and the value column is called `daily_revenue` — predict the next 30 days, and figure out which modeling framework fits best on your own." Every field beyond datetime_column, historical_data, horizon, and target_column is optional, but adding a few makes the output far more useful. 🎯
kind: operator
type: forecast
version: v1
spec:
  datetime_column:
    name: sales_date
    format: "%Y-%m-%d"
  historical_data:
    url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
  target_column: daily_revenue
  horizon: 30
  model: auto
  metric: smape
  generate_explanations: true
  output_directory:
    url: results

Save this as forecast.yaml inside the folder from Step 2.


📥 Step 4 — Run It and Read the Results

📌 What does the command below do?

This is the one command that actually trains everything: it loads your historical data, tries the candidate frameworks, keeps the best performer, generates a forecast, and writes every output file into the output_directory you configured (or a folder called results if you didn't set one). ⏳
ads operator run -f forecast.yaml

What lands in your output folder when it finishes:

  • forecast.csv — date, forecast value, lower bound, upper bound, per row
  • metrics.csv — the backtested error score for the winning model
  • global_explanations.csv / local_explanations.csv — covered in Step 7
  • report.html — a shareable visual summary a business planner can open directly
📌 What does the code below do?

This reads forecast.csv with pandas and prints the first 5 forecasted days next to their confidence range — the exact same file your dashboard or planning app would consume. 📅
import pandas as pd

forecast_df = pd.read_csv("results/forecast.csv")
print(forecast_df[["Date", "forecast_value", "lower_bound", "upper_bound"]].head())

Output:

Date           forecast_value    lower_bound    upper_bound
2026-07-06         48230.0          44100.0        52360.0
2026-07-07         51940.0          47500.0        56380.0
2026-07-08         49870.0          45600.0        54140.0
2026-07-09         53210.0          48800.0        57620.0
2026-07-10         55480.0          50900.0        60060.0

Notice every prediction comes with a range, not just one number — that range tells your business exactly how much uncertainty to plan around. 📊


🏆 Step 5 — Model Selection, Per Series

The model field accepts a specific framework name (prophet, arima, neuralprophet, automlx, autots) if you already know what fits your data, or auto to let the operator decide. When you're forecasting many related series at once — say, one per store — auto-selection can go a level deeper and recommend a different best-fit framework independently for each individual series, since a steady seller and a spiky, promotion-driven product rarely share the same ideal model.

📌 What does the YAML addition below do?

Adding target_category_columns tells the operator your CSV actually contains many interleaved series (one per store), not just one — it will train and forecast each store independently, in the same run. 🧩
spec:
  historical_data:
    url: oci://sales-data-bucket@your_namespace/all_stores_sales.csv
  target_category_columns:
    - store_id
  target_column: daily_revenue
  datetime_column:
    name: sales_date
  horizon: 30
  model: auto

Reading the winning framework per series from metrics.csv:

import pandas as pd

metrics_df = pd.read_csv("results/metrics.csv")
print(metrics_df[["store_id", "selected_model", "smape"]].head())
store_id   selected_model      smape
101        neuralprophet        4.12
102        arima                 6.03
103        prophet               5.44

A deep learning model won for store 101 — but a classic statistical method won for store 102. This is exactly why per-series automatic selection matters at real retail scale: no single algorithm is right for every store. 🥇


🔗 Step 6 — Adding External Regressors

Sales rarely move in isolation — promotions, prices, and even weather can be genuine drivers. The operator's additional_data source lets those known, plannable factors feed directly into the forecast.

💡 Real-World Analogy:

A single-series forecast is like predicting tomorrow's temperature only from yesterday's temperature. Adding regressors is like also checking whether a storm system is approaching — it uses more relevant information to make a smarter guess. 🌩️
📌 What does the YAML below do?

This trains the same forecast as before, but this time hands the operator a second file with an on_promotion and average_price column for every date in both the historical range and the future horizon. The model learns how these factors historically influenced sales, and factors that relationship into its future prediction. Critically, this second file must already contain values reaching across the whole forecast horizon — you have to know or credibly plan these values in advance. 🧩
spec:
  datetime_column:
    name: sales_date
    format: "%Y-%m-%d"
  historical_data:
    url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
  additional_data:
    url: oci://sales-data-bucket@your_namespace/promo_and_price_plan.csv
  target_column: daily_revenue
  horizon: 30
  model: auto
  generate_explanations: true
  output_directory:
    url: results
ads operator run -f forecast.yaml

🔍 Step 7 — Understanding Why the Forecast Looks the Way It Does

Setting generate_explanations: true produces two extra files that turn "trust me" into "here's exactly why."

📌 What does the code below do?

This reads global_explanations.csv — the file that describes which factors influenced the whole forecast the most overall — so a business planner can see why the model expects a spike or dip, not just that it expects one. There's a matching local_explanations.csv that breaks this down per individual prediction, not just in aggregate. 📋
import pandas as pd

global_df = pd.read_csv("results/global_explanations.csv")
print(global_df.sort_values("importance_score", ascending=False).head())
feature            importance_score
on_promotion              0.47
average_price             0.31
day_of_week                0.22

Promotions turn out to be the single biggest driver of revenue swings — more influential than even price changes. That's a genuinely useful business insight, not just a number. 💡


🔁 Step 8 — What-If Analysis & Model Deployment

A forecast trained once and never revisited slowly goes stale as real business conditions shift. The operator's newer What-If Analysis capability packages a trained model into the OCI Model Catalog and, optionally, deploys it as a live endpoint — so planners can test "what if we ran a bigger promotion next month?" scenarios without retraining from scratch.

📌 What does the YAML addition below do?

Adding a what_if_analysis block tells the operator to save the trained model bundle to the Model Catalog after this run. Including a model_deployment block goes one step further and has ADS automatically stand up a live HTTP endpoint, complete with logging and autoscaling — so a planning app can POST a hypothetical additional_data scenario and get a new forecast back instantly. 🗓️
spec:
  # ...same fields as Step 6...
  what_if_analysis:
    model_deployment:
      display_name: "sales-forecast-whatif-endpoint"
      auto_scaling:
        min_instances: 1
        max_instances: 3

What you'll find after the run:

results/deployment_info.json
  → model catalog OCID
  → model deployment OCID
  → live endpoint URL

From here, updating the promotion or price assumptions in a fresh additional_data file and POSTing it to that endpoint returns an updated forecast in real time — without a full retrain. 🚀


🧵 Step 9 — The Complete End-to-End Example, Start to Finish

Let's stitch everything together into one real, runnable spec plus the two commands that take it from an empty folder to a printed forecast table. 🧵

📌 What does this end-to-end flow do, from top to bottom?

  1. Scaffolds a new forecast project folder
  2. Points at your historical sales data already sitting in Object Storage
  3. Requests a 30-day forecast with automatic model selection and explanations
  4. Runs the operator, which trains, compares, and picks the best model
  5. Reads back the first 5 days of the forecast, with confidence ranges
Replace the placeholder bucket and namespace values with your own, then run the two CLI commands from inside the project folder. 🚀
# ──── 1. Scaffold the project ───────────────────────────────────────────────
$ ads operator init --type forecast --output daily-sales-forecast
$ cd daily-sales-forecast

# ──── 2. Edit forecast.yaml ─────────────────────────────────────────────────
cat > forecast.yaml <<'EOF'
kind: operator
type: forecast
version: v1
spec:
  datetime_column:
    name: sales_date
    format: "%Y-%m-%d"
  historical_data:
    url: oci://sales-data-bucket@your_namespace/daily_sales_history.csv
  target_column: daily_revenue
  horizon: 30
  model: auto
  metric: smape
  generate_explanations: true
  output_directory:
    url: results
EOF

# ──── 3. Run the forecast ───────────────────────────────────────────────────
$ ads operator run -f forecast.yaml
# ──── 4. Read the results ───────────────────────────────────────────────────
import pandas as pd

df = pd.read_csv("results/forecast.csv")
print(df[["Date", "forecast_value", "lower_bound", "upper_bound"]].head())

Output:

Date           forecast_value    lower_bound    upper_bound
2026-07-06         48230.0          44100.0        52360.0
2026-07-07         51940.0          47500.0        56380.0
2026-07-08         49870.0          45600.0        54140.0
2026-07-09         53210.0          48800.0        57620.0
2026-07-10         55480.0          50900.0        60060.0
✅ What You Just Built:

A complete, working pipeline that goes from raw historical sales data, to a YAML-configured forecast run, to a 30-day forecast with confidence ranges — in one short spec plus two CLI commands, using the real ADS Forecasting Operator. 🏁

🏛️ Enterprise Architecture — Where This Fits

  • 🗄️ Object Storage / Data Lake — holds historical time series data
  • 🔄 OCI Data Integration — moves and refreshes historical data on a schedule
  • 📈 OCI Data Science + ADS Forecasting Operator — trains, selects the best model per series, and produces forecasts
  • 📦 Model Catalog & Model Deployment — stores trained models and optionally serves them for live What-If scenarios
  • 📊 Oracle Analytics Cloud — visualizes forecasts and confidence ranges for business users
  • 🗂️ Business Applications — supply chain, inventory, or workforce planning tools consuming the forecast output
[ Historical Data In Object Storage ] -> [ ADS Forecasting Operator: Train + Select Best Model ]
        -> [ forecast.csv + Confidence Interval + Explanations ]
        -> [ Model Catalog (+ optional live Deployment) ]
        -> [ Oracle Analytics Cloud Dashboard ]
        -> [ Supply Chain / Planning App Consumes Forecast Output ]

🏢 A Named Enterprise Scenario: "BrightMart Retail"

Let's make this concrete. BrightMart Retail (a fictional example) runs 120 stores and previously forecasted demand using a single company-wide spreadsheet model, updated manually once a month by a small planning team.

  • Stockouts during promotional weekends were common, since the old model didn't account for promotions
  • Overstock of slow-moving items tied up warehouse space and cash
  • The planning team could only realistically update forecasts for the 20 highest-revenue stores

After adopting the ADS Forecasting Operator with a per-store, promotion-aware spec:

  • All 120 stores now get their own forecast in a single run, using target_category_columns to keep every store's series independent
  • Promotional periods are factored in directly, since planned promotion and price data feeds in through additional_data
  • Stockouts during promotions dropped significantly, since inventory could be pre-positioned ahead of expected demand spikes
  • The planning team now spends its time reviewing global_explanations.csv and exceptions, not manually rebuilding spreadsheets
  • A live What-If deployment lets regional managers test "what if we extend the promotion a week?" without waiting for a full retrain
✅ The Real Lesson:

The win wasn't a smarter single model — it was being able to give every single store its own properly-tuned, automatically-refreshed forecast, something that was simply impossible for a small human team to do manually at that scale. 📈

🏆 Best Practices

  • 📅 Feed in as much clean history as you have — more history generally helps the model learn seasonality correctly
  • 🔗 Add known external factors via additional_data — promotions, price changes, and holidays often meaningfully improve accuracy
  • 🎯 Always check the confidence interval, not just the single forecast number, when making planning decisions
  • 🥇 Review metrics.csv per series — a close second-place model can be a useful sanity check, especially with model: auto
  • 🔍 Use global_explanations.csv and local_explanations.csv to validate that the model's reasoning actually matches real business knowledge
  • 🔁 Re-run on a schedule — stale forecasts quietly become inaccurate as real conditions shift
❌ Common Mistakes to Avoid:

  • Do NOT feed in very short histories (a few weeks) and expect the model to learn seasonality
  • Do NOT ignore the confidence interval — planning around only the midpoint forecast hides real risk
  • Do NOT put a factor in additional_data if you won't actually know its future value in advance — the file must cover the full horizon, not just history
  • Do NOT disable default preprocessing steps (missing-value imputation, outlier treatment) without understanding why — it can cause a framework to fail outright
  • Do NOT forget to re-check forecasts after a major business disruption — the historical pattern the model learned may no longer apply

🌍 Real-World Use Cases

  • 🛒 Retail: Product demand forecasting per store, accounting for promotions and seasonality
  • 🏭 Manufacturing: Raw material and component demand planning
  • 💰 Finance: Revenue and cash flow forecasting for budgeting cycles
  • ☎️ Customer Support: Call center volume forecasting for staffing decisions
  • ⚡ Utilities & Energy: Demand forecasting influenced by weather conditions
  • 🖥️ IT Operations: Forecasting compute or storage capacity needs ahead of demand spikes

  • Per-Series Model Selection at Scale — modern forecasting workloads increasingly select the best algorithm independently for each individual series, since the best-performing framework genuinely varies by frequency, horizon, and signal pattern
  • Real-Time What-If Analysis — trained models packaged into the Model Catalog and optionally deployed live, so planners can test scenarios without a full retrain
  • Forecasting Feeding Generative AI Narratives — explanation outputs increasingly feed directly into Generative AI, which drafts plain-language planning summaries for business stakeholders
  • Explainability As Standard Practice — global and local influencer breakdowns are increasingly treated as a required output, not an optional extra, especially in regulated planning processes

📝 Quick Summary — What We Learned

  • What OCI Forecasting is → The ADS Forecasting Operator, a low-code, YAML-driven capability inside OCI Data Science that delivers time-series forecasts via statistical, ML, and deep learning frameworks, with no data science expertise required
  • Automated Model Selection → Set model: auto to try multiple frameworks and keep the one with the strongest backtested accuracy, per series if needed
  • Single-Series vs. Regressor-Enriched → Use only the metric's own history, or enrich it with additional_data like promotions and price
  • Confidence Intervals → Every forecast comes with a range, not just a single guessed number
  • Explainability → global_explanations.csv and local_explanations.csv show exactly why the model predicted what it did
  • What-If Analysis → Package and optionally deploy a model for live scenario testing, not just static reports
  • Enterprise Pattern → Scale properly-tuned, explainable forecasts across every product, store, or metric that matters — not just the few a small team could handle manually

❓ Troubleshooting & Frequently Asked Questions

How much historical data do I actually need?

There's no hard minimum, but generally at least 2-3 full seasonal cycles (often 2 years for daily/weekly business data) gives the model enough signal to properly learn recurring patterns.

Can I forecast several related metrics at once, like sales for many products?

Yes — add the relevant column(s) to target_category_columns so the operator trains and forecasts a series per group, such as one per product or one per store, within the same run.

What if my forecast run fails?

The most common cause is a mismatch between datetime_column's name/format and what's actually in your file, or an additional_data file that doesn't extend across the full forecast horizon. Check those two first before anything else.

Does this replace a data science team entirely?

It removes the need to hand-build classical forecasting models for routine business metrics. Highly specialized, novel forecasting problems may still benefit from custom modeling directly in OCI Data Science, but most enterprise demand and revenue forecasting fits comfortably within this operator.

Is "OCI Forecasting" the same thing as the 2021 announced service?

Conceptually yes — same goal, same core ideas (auto model selection, confidence intervals, explainability). The 2021 announcement described a limited-availability standalone API; the generally available path today is the ADS Forecasting Operator inside OCI Data Science, which is what this guide covers.

Happy building! 📈✨

Comments