Imagine your company gets 10,000 messages every single day — emails, support tickets, reviews, and chats, written in a dozen different languages. Reading all of that by hand is impossible! OCI Language is like hiring a tireless robot reader who can open every single message, understand what it means, and hand you a clean, structured answer in seconds.
🧠 What is OCI Language?
OCI Language is an AI service on Oracle Cloud Infrastructure (OCI) that reads plain text and instantly tells you what it means — without you writing any machine learning code, and without needing a data science degree.
Think of it like this 👇
- A customer writes: "Terrible experience, my order from Acme Corp never arrived!" ✉️
- You send that sentence to OCI Language
- It tells you: "Sentiment: Negative. Entity found: Acme Corp (Organization)."
- No human reading required! ✅
Think of OCI Language as a super-fast librarian who has read every book in the world. Hand them any sentence, in almost any language, and they instantly tell you the mood, the important names, and even hide anything private — all without ever getting tired. 📚
🌟 What Can OCI Language Do?
OCI Language has several superpowers hiding inside one toolbox. Let's meet them all:
- 🌐 Language Detection — Figures out which of 75+ languages a piece of text is written in
- 🗂️ Text Classification — Sorts documents into categories, like "Sports" or "Finance"
- 🏷️ Named Entity Recognition (NER) — Finds names, places, companies, dates, and emails hidden in text
- 🔑 Key Phrase Extraction — Pulls out the most important short phrases from a long text
- 😊 Sentiment Analysis — Tells you if a message is positive, negative, or neutral — even topic by topic
- 🕵️ PII Detection & Masking — Finds sensitive data like card numbers and phone numbers and hides them
- 🔄 Text Translation — Converts text from one language into another
- 🎓 Custom Model Training — Teach the AI your own company-specific words and categories
OCI Language is a serverless, multi-tenant service — you don't manage any servers. You simply call an API, and Oracle's pretrained models do all the heavy thinking. 🧩
OCI Language does not write new stories, essays, or long creative replies. That job belongs to OCI Generative AI. OCI Language reads and understands — Generative AI writes and creates. We'll show how they team up later in this blog. ✍️
🗺️ Big Picture — How Does It All Work?
Before we touch any code, let's follow one message on its full journey through OCI:
┌───────────────┐ ┌──────────────────────┐ ┌─────────────────────────┐
│ Your Text │────►│ IAM Security Gate │────►│ OCI Language Service │
│ (email/chat) │ │ (checks who you are)│ │ (the reading brain) │
└───────────────┘ └──────────────────────┘ └────────────┬────────────┘
│
┌────────────▼────────────┐
│ JSON Result │
│ Sentiment · Entities · │
│ Language · Categories │
└─────────────────────────┘
Simple flow: Send text → AI reads it → Get structured answers back. 🎯
🏗️ Step 1 — Set Up Your OCI Environment
Before we call OCI Language, we need to prepare our tools. Think of this like sharpening your pencils before an exam! ✏️
Step 1a: Create an OCI Account
- Go to cloud.oracle.com
- Click "Start for free"
- Enter your name, email, and set a password
- Verify your email ✅
- Done! You now have access to Oracle Cloud
Step 1b: Install the OCI Python SDK
It tells your computer: "Please download the OCI toolbox so my Python code can talk to Oracle Cloud." You only need to run this once.
pip install oci
Step 1c: Configure Your Credentials
OCI needs to know who you are before it lets your code in — like showing your ID card at a security gate. 🪪
It starts a small setup wizard, asks you a few questions (your region, your user ID), and creates a private key file on your computer. OCI checks this key every single time your code makes a request. 🔑
oci setup config
Follow the prompts. It creates a config file at ~/.oci/config. That's it! ✅
Your OCI administrator needs to add one line of policy so your user group is allowed to call the Language service. Think of it as adding your name to a guest list. 📋
allow group MyDevelopers to use ai-service-language-family in tenancy
Never share your
~/.oci/config file or private key with anyone.
It's like sharing your bank PIN — keep it completely secret! 🔒
🌐 Step 2 — Detecting Which Language a Sentence Is Written In
Before you can understand a message, you first need to know which language it is written in. This tool reads a sentence and guesses the language, along with a confidence score that tells you how sure it is. 🎯
This code sets up a connection to OCI Language, wraps one sentence inside a small document object, and asks the AI: "What language is this?" The AI reads it and hands back a language code, such as "en" for English. 🌍
import oci
# Load your OCI configuration (region, key path, tenancy) from the config file
config = oci.config.from_file()
# Create the client object that talks to OCI Language
ai_client = oci.ai_language.AIServiceLanguageClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
# Step 1: wrap the text inside a document object with a unique key
doc = oci.ai_language.models.DominantLanguageDocument(
key="doc1",
text="Bonjour, comment allez-vous aujourd'hui?"
)
# Step 2: put the document inside a batch request
details = oci.ai_language.models.BatchDetectDominantLanguageDetails(
documents=[doc],
compartment_id=compartment_id
)
# Step 3: send the batch to OCI Language and read the answer
response = ai_client.batch_detect_dominant_language(details)
top_language = response.data.documents[0].languages[0]
print(f"Detected Language : {top_language.code}")
print(f"Confidence Score : {top_language.score * 100:.1f}%")
Output:
Detected Language : fr Confidence Score : 99.4%
The AI is 99.4% sure that sentence is French! 🇫🇷
🗂️ Step 3 — Text Classification
Imagine a librarian who reads any document and instantly knows which shelf it belongs on. Text Classification does exactly that for your text, sorting it into categories like "Finance", "Sports", or "Technology". 📚
This code sends a news snippet to OCI Language and asks: "What topic is this about?" The AI reads it and returns one or more category labels, along with a confidence score for each. 🏷️
text_doc = oci.ai_language.models.TextDocument(
key="news1",
text="The central bank raised interest rates by half a percent today.",
language_code="en"
)
classify_details = oci.ai_language.models.BatchDetectLanguageTextClassificationDetails(
documents=[text_doc],
compartment_id=compartment_id
)
classify_response = ai_client.batch_detect_language_text_classification(classify_details)
for label in classify_response.data.documents[0].text_classification:
print(f"Category: {label.label:15} Confidence: {label.score * 100:.1f}%")
Output:
Category: FINANCE Confidence: 96.2% Category: ECONOMY Confidence: 88.7%
Now this article can be automatically filed under the right topic — no manual tagging needed! ✅
🏷️ Step 4 — Named Entity Recognition (NER)
This tool works like a smart highlighter, marking out the important words in a sentence — names of people, companies, places, and dates. 🖍️
This code reads a sentence and pulls out every important entity it can find, telling you not just the word, but also what type of thing it is — a person, an organization, or a place. 🔍
entity_doc = oci.ai_language.models.TextDocument(
key="msg1",
text="Sarah from Acme Corp visited Mumbai last Tuesday.",
language_code="en"
)
entity_details = oci.ai_language.models.BatchDetectLanguageEntitiesDetails(
documents=[entity_doc],
compartment_id=compartment_id
)
entity_response = ai_client.batch_detect_language_entities(entity_details)
for entity in entity_response.data.documents[0].entities:
print(f"{entity.text:15} → {entity.type}")
Output:
Sarah → PERSON Acme Corp → ORGANIZATION Mumbai → LOCATION last Tuesday → DATE
Four important pieces of information, pulled out automatically in one API call! 🎯
🔑 Step 5 — Key Phrase Extraction
Sometimes you don't need every word — you just want the highlights. Key Phrase Extraction reads a big block of text and hands you back only the most important short phrases, like a highlighter that skips the boring parts. ✨
This code sends a paragraph to OCI Language and gets back a short list of the most meaningful phrases — perfect for building quick summaries or tag clouds. ☁️
phrase_doc = oci.ai_language.models.TextDocument(
key="review1",
text="The new laptop has a stunning display, long battery life, and a lightweight design.",
language_code="en"
)
phrase_details = oci.ai_language.models.BatchDetectLanguageKeyPhrasesDetails(
documents=[phrase_doc],
compartment_id=compartment_id
)
phrase_response = ai_client.batch_detect_language_key_phrases(phrase_details)
for phrase in phrase_response.data.documents[0].key_phrases:
print(f"- {phrase.text}")
Output:
- stunning display - long battery life - lightweight design
A perfect three-word summary of a much longer review! 📝
😊 Step 6 — Sentiment Analysis (Including Aspect-Level)
This tool reads a message and tells you the mood — positive, negative, or neutral. The advanced version can even judge the mood about one specific part of a sentence, separately from the rest. 🎭
This code reads a restaurant review and finds the overall mood of the whole sentence, and also the mood about each separate topic mentioned inside it — like food versus service. 🍽️
sentiment_doc = oci.ai_language.models.TextDocument(
key="review2",
text="The food was great but the service was slow.",
language_code="en"
)
sentiment_details = oci.ai_language.models.BatchDetectLanguageSentimentsDetails(
documents=[sentiment_doc],
compartment_id=compartment_id
)
# level=["ASPECT","SENTENCE"] asks for both whole-sentence and per-topic sentiment
sentiment_response = ai_client.batch_detect_language_sentiments(
sentiment_details,
level=["ASPECT", "SENTENCE"]
)
doc_result = sentiment_response.data.documents[0]
print(f"Overall Sentence Sentiment: {doc_result.sentences[0].sentiment}")
for aspect in doc_result.aspects:
print(f" Aspect: {aspect.text:10} → {aspect.sentiment}")
Output:
Overall Sentence Sentiment: Mixed Aspect: food → Positive Aspect: service → Negative
Now you know exactly which part of the experience needs improvement, instead of just a vague overall score. 🎯
🕵️ Step 7 — Detecting and Masking Sensitive Information (PII)
This tool scans text for sensitive details — credit card numbers, phone numbers, birth dates — and hides them with symbols, so private data never leaks into your logs or dashboards. 🔒
This code scans a customer message for a credit card number, then replaces most of the digits with asterisks — keeping only the last 2 digits visible for reference. 🎭
pii_doc = oci.ai_language.models.TextDocument(
key="pii1",
text="My card number is 4111111111111234, please refund it.",
language_code="en"
)
masking_rule = oci.ai_language.models.PiiEntityMask(
mode="MASK",
masking_character="*",
leave_characters_unmasked=2,
is_unmasked_from_end=True
)
pii_details = oci.ai_language.models.BatchDetectLanguagePiiEntitiesDetails(
documents=[pii_doc],
compartment_id=compartment_id,
masking={"CREDIT_DEBIT_NUMBER": masking_rule}
)
pii_response = ai_client.batch_detect_language_pii_entities(pii_details)
print(pii_response.data.documents[0].masked_text)
Output:
My card number is ************************34, please refund it.
Banks, hospitals, and insurance companies are legally required to protect private data. Masking it right at the extraction stage means unsafe data never even reaches your database. 🏦
🔄 Step 8 — Translating Text Into Another Language
Global companies get messages in every language imaginable. Text Translation converts a sentence from one language into another, so a single support team can understand customers from anywhere in the world. 🌏
This code takes a French sentence and asks OCI Language to rewrite it in English, so your support agent can read and reply without knowing a word of French. 🗣️
translate_doc = oci.ai_language.models.TextDocument(
key="t1",
text="Merci pour votre aide rapide.",
language_code="fr"
)
translate_details = oci.ai_language.models.BatchLanguageTranslationDetails(
documents=[translate_doc],
compartment_id=compartment_id,
target_language_code="en"
)
translate_response = ai_client.batch_language_translation(translate_details)
print(translate_response.data.documents[0].translated_text)
Output:
Thank you for your quick help.
🏭 Step 9 — Building a Real-World Text Pipeline
Let's now combine everything into one real scenario. Imagine you run a support desk that receives hundreds of messages every day. You want each message automatically read, scored for mood, cleaned of private data, and saved into a CSV report. 📑➡️📊
This is a complete automation script that:
- Reads a list of incoming customer messages
- Detects the sentiment of each message
- Masks any sensitive numbers found inside
- Saves everything into a clean CSV report
import oci
import csv
config = oci.config.from_file()
ai_client = oci.ai_language.AIServiceLanguageClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
incoming_messages = [
"My card 4111111111111234 was charged twice, I am very upset!",
"Thanks for the fast delivery, everything arrived perfectly.",
"The app keeps crashing, please call me at 987-654-3210."
]
masking_rule = oci.ai_language.models.PiiEntityMask(
mode="MASK", masking_character="*",
leave_characters_unmasked=2, is_unmasked_from_end=True
)
with open("support_report.csv", "w", newline="") as csvfile:
writer = csv.writer(csvfile)
writer.writerow(["Clean Message", "Sentiment"])
for i, message in enumerate(incoming_messages):
doc = oci.ai_language.models.TextDocument(
key=f"msg{i}", text=message, language_code="en"
)
# detect sentiment for the whole message
sentiment_details = oci.ai_language.models.BatchDetectLanguageSentimentsDetails(
documents=[doc], compartment_id=compartment_id
)
sentiment_result = ai_client.batch_detect_language_sentiments(
sentiment_details, level=["SENTENCE"]
)
mood = sentiment_result.data.documents[0].sentences[0].sentiment
# mask any PII found in the same message
pii_details = oci.ai_language.models.BatchDetectLanguagePiiEntitiesDetails(
documents=[doc], compartment_id=compartment_id,
masking={"CREDIT_DEBIT_NUMBER": masking_rule, "PHONE_NUMBER": masking_rule}
)
pii_result = ai_client.batch_detect_language_pii_entities(pii_details)
clean_text = pii_result.data.documents[0].masked_text
writer.writerow([clean_text, mood])
print(f"✅ Processed message {i + 1}: {mood}")
print("\n🏁 Done! Report saved to 'support_report.csv'")
Output:
✅ Processed message 1: Negative ✅ Processed message 2: Positive ✅ Processed message 3: Negative 🏁 Done! Report saved to 'support_report.csv'
What used to take a whole support team hours now runs automatically in seconds! 💪
🎓 Step 10 — Training a Custom Language Model
The pretrained models are great for general text, but every company has its own special words — product codes, internal team names, or industry jargon. Custom Model Training lets you teach OCI Language your own vocabulary. 🧠
Training a custom model is like teaching a new employee your company's internal shorthand. You show them 50–100 labelled examples, and after a short training session, they understand your special terms perfectly. 👩🏫
The Training Steps (High Level)
- Step A: Collect sample text records — support tickets, product descriptions, etc.
- Step B: Use OCI Data Labeling to mark the categories or entities in each sample
- Step C: Create a custom model project in OCI Language and point it to your labelled data
- Step D: Let OCI train the model automatically
- Step E: Deploy the trained model to a dedicated endpoint and start calling it, same as any built-in model
You can assign more or fewer "inference units" (compute power) to your custom model's endpoint depending on how much traffic you expect — just like choosing a bigger or smaller engine for a car. 🚗
🏛️ Enterprise Architecture — Where OCI Language Fits
A real enterprise system is like a big kitchen with many chefs, and OCI Language is only one of them. Let's meet the whole team:
- 🚪 API Gateway — the front door that safely receives requests from outside apps
- ⚡ OCI Functions — small pieces of code that run only when needed
- 🌊 OCI Streaming / Events — a conveyor belt moving new messages continuously
- 🧠 OCI Language — the reading chef, understanding meaning, mood, and entities
- 🗄️ Autonomous Database / OCI 23ai — the pantry, storing structured results and vector data
- ✍️ OCI Generative AI Agents — the creative chef, writing replies and taking multi-step actions
[ Customer Message ] → [ API Gateway ] → [ OCI Functions ]
→ [ OCI Language: Sentiment + NER + PII Mask ]
→ [ Autonomous DB / 23ai Vector Store ]
→ [ OCI Generative AI Agent drafts a smart reply ]
→ [ Human Reviewer Approves ] → [ Reply Sent ]
This pattern is popular in 2026 because it blends two worlds: OCI Language gives fast, cheap, structured understanding, while OCI Generative AI Agents add flexible reasoning and natural-sounding replies — only after the cheaper step has already filtered the noise. ⚙️
🏆 Best Practices
- 📦 Batch your requests — send multiple documents in one API call instead of one at a time; it's faster and cheaper
- ✂️ Split very long documents — break huge text blocks into smaller chunks for better accuracy
- 🔐 Always mask PII before storage — never let unmasked sensitive data touch your database or logs
- 🎯 Set confidence thresholds — only trust results above a certain confidence score, such as 80%
- 🎓 Train custom models for jargon — when your industry has words the general model doesn't know
- 🔑 Use OCI Vault for secrets — never hardcode compartment IDs or keys directly in your code
- Do NOT skip PII masking to save a little processing time — it creates serious legal risk
- Do NOT send one document per API call when you have hundreds waiting — always batch them
- Do NOT expect OCI Language to generate long creative text — pair it with OCI Generative AI for that
- Do NOT ignore confidence scores — a low score means the AI itself isn't fully sure
🌍 Real-World Use Cases
- 🏦 Banking & Finance: Detect fraud-related sentiment and mask card numbers in support chats
- 🏥 Healthcare: Hide patient names and birth dates in shared documents automatically
- 🛒 Retail & E-Commerce: Track sentiment across product reviews in dozens of languages
- ✈️ Travel & Airlines: Auto-route complaint tickets by topic and urgency
- 📰 Media & Publishing: Auto-tag articles by category for faster organization
- 🎧 Global Support Desks: Translate and analyze incoming messages from any country in real time
🚀 What's New and Trending in 2026
- Pairing With Generative AI Agents — using OCI Language as a fast, cheap filter before a heavier generative agent responds
- Vector Search With OCI 23ai — storing extracted entities and key phrases alongside vector embeddings for smarter search and retrieval-augmented generation
- Privacy-By-Default Pipelines — PII masking wired directly into ingestion, not added as an afterthought
- Event-Driven Text Processing — triggering OCI Language instantly through OCI Events or Streaming the moment new text arrives, instead of running batch jobs once a day
📝 Quick Summary — What We Learned
- What OCI Language is → A serverless AI service that reads and understands plain text
- Language Detection → Figures out which language a sentence is written in
- Text Classification → Sorts text into categories automatically
- Named Entity Recognition → Finds names, places, organizations, and dates
- Key Phrase Extraction → Pulls out the most important short phrases
- Sentiment Analysis → Detects mood, even topic by topic
- PII Detection & Masking → Hides sensitive data automatically
- Text Translation → Converts text between languages
- Custom Models → Teach the AI your own company-specific vocabulary
- Enterprise Pattern → Combine OCI Language with Generative AI Agents for fast, cost-efficient, smart pipelines
Happy building! 🧠✨
Comments
Post a Comment