OCI State-of-the-Art NLP Models — The World's Best AI Brains, Available on Your Cloud
Imagine you could hire the world's smartest language expert — someone who has read every book, article, and website ever written. They can summarise a 200-page report in 3 bullet points, answer any question about your company's documents, write professional emails in 12 languages, generate code from plain English instructions, and have a natural, intelligent conversation about anything.
That super-expert is now available as a cloud API — and Oracle Cloud Infrastructure (OCI)
hosts the most powerful collection of these AI language brains available anywhere.
They are called State-of-the-Art (SOTA) NLP Models,
and you can start using them today with a few lines of Python code!
What Does "State-of-the-Art NLP Models" Mean?
Breaking It Down Word by Word
- 📖 NLP = Natural Language Processing. "Natural Language" means human language — the kind you use when texting, speaking, or writing emails. "Processing" means a computer understanding and working with it. So NLP = computers understanding human language!
- 🏆 State-of-the-Art (SOTA) = The absolute best currently available. When someone says "SOTA", they mean: "This is the most advanced, most capable version of this technology that exists today." SOTA NLP models are the AI models that score highest on benchmarks — the top performers across every language task.
- 🧠 Models = Think of a model as a highly trained brain. Just like a chess grandmaster's brain is trained on millions of games, an NLP model is trained on hundreds of billions of words of text. After training, it can do remarkable things with language.
💡 The simple version: SOTA NLP Models are the world's smartest AI language brains. OCI gives you access to them through a simple API — like renting brain power from the best AI labs in the world, without building or maintaining anything yourself!
What Are Large Language Models (LLMs)? 📚
The most powerful NLP models are called Large Language Models (LLMs). They are "large" because they have billions of parameters (the internal settings the model learned during training). They are "language models" because they were trained to understand and generate human language.
Here's an analogy that makes LLMs easy to understand:
HOW AN LLM WORKS — The "Autocomplete on Steroids" Model: Your phone's keyboard has autocomplete: You type "Happy Birth..." → it suggests "day" ✅ Simple. Limited. Only knows your recent messages. An LLM does the same thing — but at a superhuman level: Input: "Explain quantum entanglement to a 10-year-old" LLM sees this → predicts the best possible next word, then the next word, then the next — billions of times — building up a complete, accurate, beautifully written explanation. But LLMs aren't just autocomplete: They've read billions of pages of human knowledge during training. They understand context, reasoning, cause and effect, and nuance. They can solve maths problems, write code, translate languages, summarise documents, answer questions, and hold conversations.
The "size" of a model is measured in parameters — the internal numbers the model learned. More parameters (usually) = smarter model, but also more expensive to run.
MODEL SIZE COMPARISON: Small model: ~7 billion parameters → Fast, cheap, good for simple tasks Medium model: ~70 billion parameters → Balanced, good for most enterprise tasks Large model: ~400 billion parameters → Slow, expensive, best for complex reasoning
Think of parameters like the dials on a giant soundboard mixer — each one is a tiny setting that was tuned during training. A 70 billion-parameter model has 70 billion such dials, all carefully set to make the model smart. During training, the model read massive amounts of text and adjusted all those dials millions of times until it could predict language accurately. That's how it "learned"!
What SOTA NLP Models Are Available on OCI?
OCI Generative AI is a fully managed service — Oracle hosts these giant models on powerful GPU clusters so you don't have to. You just call the API and get results. OCI offers models from five leading AI labs:
OCI GENERATIVE AI MODEL CATALOGUE ═══════════════════════════════════════════════════════════════════ 🔵 COHERE MODELS ├── Command A (03-2025) ─── Context: 256K tokens ─── Best: RAG, Agents, Multilingual ├── Command A Vision ─── Context: 256K tokens ─── Best: Image + text tasks ├── Command A Reasoning ─── Context: 256K tokens ─── Best: Step-by-step reasoning ├── Command R+ ─── Context: 128K tokens ─── Best: Complex Q&A, Summarisation ├── Command R ─── Context: 128K tokens ─── Best: RAG, Enterprise workflows ├── Embed v4.0 ─── Multimodal embeddings ─── Best: Semantic search └── Embed English/Multilingual v3.0 ──────────────── Best: Text vector search 🦙 META LLAMA MODELS ├── Llama 4 Maverick (17B active / ~400B total) ── Best: Complex reasoning, coding ├── Llama 4 Scout (17B active / ~109B total) ── Best: Cost-efficient tasks └── Llama 3.3 70B Instruct ───────────────────── Best: Instruction following, code 🔷 GOOGLE GEMINI MODELS (via OCI partnership) └── Gemini 2.5 Pro / Flash / Flash-Lite ────── Best: Multimodal, long context ⚡ xAI GROK MODELS └── Grok 4 / Grok 3 Mini Fast ──────────────── Best: Real-time reasoning tasks 🤖 OPENAI GPT MODELS (gpt-oss variants) └── GPT-4.1 and variants ────────────────────── Best: General-purpose tasks ═══════════════════════════════════════════════════════════════════ All accessible via: https://inference.generativeai.us-chicago-1.oci.oraclecloud.com
Key NLP Tasks These Models Can Perform 🎯
Before we write code, let's understand the main tasks that SOTA NLP models excel at. Think of these as the different "superpowers" your AI language brain has:
-
💬 Chat / Conversational AI
Have a natural, multi-turn conversation. The model remembers the entire conversation history.
Example: "Build me a customer support chatbot that answers questions about our product warranty." -
📋 Text Summarisation
Take a long document and produce a concise summary.
Example: Condense a 50-page legal contract to a 5-bullet executive summary. -
✍️ Text Generation / Content Creation
Generate emails, reports, blog posts, product descriptions, code, and more from a short prompt.
Example: "Write a professional email declining a meeting request politely." -
🔍 Question Answering (Grounded / RAG)
Answer questions based on YOUR specific documents — not just the model's training data.
Example: "Answer employee questions using our internal HR policy documents." -
🌐 Translation
Translate text between dozens of languages with high accuracy and natural fluency.
Example: "Translate our support knowledge base into French, German, Spanish, and Japanese." -
💻 Code Generation
Write code from plain English descriptions, explain code, or fix bugs.
Example: "Write a Python function that reads a CSV and creates a bar chart." -
🔢 Text Embeddings (Semantic Search)
Convert text into mathematical vectors that capture meaning.
Two sentences with similar meaning will have similar vectors — even if the exact words differ.
Example: "Find documents semantically related to this query — even if keywords don't match."
Setting Up Your Environment 🛠️
Step 1: Install Required Libraries
These install three Python packages:
oci— Oracle's official SDK to communicate with OCI serviceslangchain-oci— The official LangChain integration for OCI (as of 2025, this replaces the older community integration)faiss-cpu— A fast library for searching through vectors (used in semantic search examples)
pip install oci pip install langchain-oci pip install faiss-cpu
Step 2: Set Up OCI Config File
The OCI config file is your ID card for Oracle Cloud. The Python SDK reads this file to verify who you are before allowing API calls. Create an API key in OCI Console → My Profile → API Keys, download the private key, and create the file at
~/.oci/config with your own values.
[DEFAULT] user=ocid1.user.oc1..aaaaaaaa...your-user-ocid... fingerprint=aa:bb:cc:dd:ee:ff:00:11:22:33:44:55:66:77:88:99 tenancy=ocid1.tenancy.oc1..aaaaaaaa...your-tenancy-ocid... region=us-chicago-1 key_file=~/.oci/oci_api_key.pem
https://inference.generativeai.us-chicago-1.oci.oraclecloud.com
(US Midwest, Chicago). Set your config region to us-chicago-1 for best model availability.
Step 3: Set the IAM Policy
An OCI administrator must grant your user group permission to use the Generative AI service. Without this, every API call fails. Add this as a Policy statement in OCI Console → Identity → Policies.
allow group <your-group-name> to manage generative-ai-family in tenancy
Use Case 1: Chat with a SOTA LLM — Your First Conversation 💬
What Is a Chat Interaction?
Chat is the most fundamental interaction with a SOTA model. You send a message (called a prompt) and the model sends back a response. The model maintains the entire conversation history so it can give contextually correct follow-up answers — just like texting a very smart friend who remembers everything you've said in the conversation.
This code connects to OCI Generative AI and sends a question to the Meta Llama 4 Maverick model — one of the most powerful models available on OCI .The model reads your question and writes back a complete, intelligent answer. Think of it as sending a text message to the world's smartest AI assistant and getting an instant reply! We then run a second question to show the model can handle a follow-up in the same conversation.
from langchain_oci import ChatOCIGenAI
from langchain_core.messages import HumanMessage, SystemMessage
# ── STEP 1: Define the OCI Generative AI endpoint ────────────────────────────
OCI_ENDPOINT = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com"
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
# ── STEP 2: Connect to OCI and choose your model ─────────────────────────────
# Llama 4 Maverick is Meta's most powerful model available on OCI.
# It uses a Mixture-of-Experts architecture: 17B active parameters from ~400B total.
# This makes it incredibly powerful but still fast to use.
llm = ChatOCIGenAI(
model_id="meta.llama-4-maverick-17b-128e-instruct-fp8",
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={
"temperature": 0.7, # 0 = very precise/deterministic, 1 = creative/varied
"max_tokens": 1000, # Maximum words in the response
"top_p": 0.9 # Controls diversity of word choice
}
)
# ── STEP 3: Build a conversation ──────────────────────────────────────────────
# A "SystemMessage" sets the behaviour and role of the AI
# A "HumanMessage" is the user's question
messages = [
SystemMessage(content=(
"You are a helpful data science assistant for beginners. "
"Always explain concepts using simple real-world analogies. "
"Keep answers concise and friendly."
)),
HumanMessage(content="What is machine learning? Explain it like I am 10 years old.")
]
# ── STEP 4: Send the conversation to OCI and get the response ─────────────────
print("🧠 Sending question to Meta Llama 4 Maverick on OCI...")
print("-" * 60)
response = llm.invoke(messages)
print(" AI Response:")
print(response.content)
# ── STEP 5: Continue the conversation with a follow-up ───────────────────────
print("\n" + "=" * 60)
print("💬 Following up with a second question...")
print("=" * 60)
# Add the AI's previous answer and a new follow-up question
messages.append(response) # Add AI's previous answer to history
messages.append(HumanMessage(content="Great! Now give me a real-world example of machine
learning that I use every day without knowing it."))
response2 = llm.invoke(messages)
print(" AI Follow-up Response:")
print(response2.content)
Example output:
🧠 Sending question to Meta Llama 4 Maverick on OCI... ------------------------------------------------------------ AI Response: Imagine you have a dog. 🐕 Every morning, you point at a bird and say "bird!" and point at a cat and say "cat!" After seeing 1,000 examples, your dog starts to recognise birds and cats on its own — without you pointing anymore. Machine learning works the same way! - You show the computer thousands of examples (like pictures of cats and dogs) - You tell it the correct answer each time - The computer learns the patterns - Then it can recognise new examples it has NEVER seen before — all by itself! The "learning" part happens automatically, from data — not from a human writing rules like "cats have pointy ears". The computer figures out the rules itself. ============================================================ 💬 Following up with a second question... ============================================================ AI Follow-up Response: You use machine learning every single day — probably dozens of times! 🎵 Spotify — When Spotify recommends a new song you end up loving, that's ML. It learned your music taste from millions of plays and skips. 📧 Gmail Spam Filter — When Gmail automatically catches spam emails and puts them in the Junk folder, that's ML. It learned what spam looks like from billions of emails. 📱 Face ID — When your phone recognises your face and unlocks instantly, that's ML. It learned the unique geometry of your face from thousands of training examples. 🎬 Netflix Recommendations — "Because you watched X, you might like Y" — that's ML finding patterns in what millions of users watched and enjoyed. Every time an app seems to "just know" what you want, ML is behind it!
Two intelligent, contextually connected responses — powered by one of the world's most advanced AI models, running on OCI! 🚀
Understanding Key Parameters — Temperature, Top-P, Max Tokens 🎛️
When you call a SOTA model, you can control HOW it responds using parameters. Think of these like the settings on a musical instrument — they change the character of the output.
-
🌡️ Temperature (0.0 to 1.0)
Controls how creative vs predictable the model is.temperature=0.0→ Always picks the most likely next word. Very precise, deterministic. Good for: facts, code, summaries.temperature=0.7→ Mixes likely and unlikely word choices. Natural and varied. Good for: general chat, explanations.temperature=1.0→ Very random, very creative. Sometimes brilliant, sometimes strange. Good for: creative writing, brainstorming.
💡 Think of temperature like a jazz musician: 0.0 = always plays the sheet music exactly. 1.0 = full improvisation! -
📊 Top-P (0.0 to 1.0, also called Nucleus Sampling)
Controls the pool of words the model considers at each step.top_p=0.9means: "Only consider words that together account for 90% of the probability." Lower = more focused, higher = more diverse.
Used together with temperature to fine-tune response diversity. -
📏 Max Tokens
Sets the maximum length of the response. One "token" ≈ 0.75 words.max_tokens=100→ Very short response.max_tokens=2000→ Long, detailed response.
More tokens = slightly higher cost (you pay per token in OCI Generative AI).
Use Case 2: Text Summarisation — Condense Long Documents Instantly 📋
Imagine your boss drops a 5-page earnings report on your desk and says "summarise this in 3 bullet points — I have 2 minutes." This code sends that long document to the Cohere Command A model and asks it to produce three clear, concise bullet-point takeaways. We use
temperature=0.1 (very low) because for summaries we want precision, not creativity!
from langchain_oci import ChatOCIGenAI
from langchain_core.messages import HumanMessage, SystemMessage
OCI_ENDPOINT = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com"
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
# ── Use Cohere Command A — excellent at summarisation and RAG tasks ───────────
# Command A has a 256K context window, meaning it can handle VERY long documents!
summariser_llm = ChatOCIGenAI(
model_id="cohere.command-a-03-2025",
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={
"temperature": 0.1, # Low temperature: we want factual, precise summaries
"max_tokens": 500
}
)
# ── The long document to summarise ───────────────────────────────────────────
long_document = """
Oracle Corporation delivered exceptional financial results for Q3 fiscal year 2026.
Total quarterly revenue reached $16.9 billion, representing a 17% increase compared
to the same quarter last year. Cloud services and license support revenue, which forms
the largest segment, grew to $10.8 billion — up 12% year-over-year.
The Infrastructure Cloud (OCI) segment was the standout performer, posting remarkable
growth of 49% year-on-year to reach $2.7 billion. This acceleration was driven by
strong enterprise demand for OCI AI and GPU infrastructure, particularly from customers
building Generative AI applications at scale. Oracle's management highlighted that OCI
GPU bookings continued to exceed supply capacity, with a waiting list extending into
Q4 2026.
Database and cloud services revenue grew 15% to $4.1 billion, reflecting rapid adoption
of Oracle Database 23ai and its integrated vector search capabilities. The Autonomous
Database subscriber count crossed 28,000 paying customers during the quarter.
Net income for the quarter was $3.4 billion, with diluted earnings per share of $1.22.
The company repurchased $3.1 billion of its own shares during the period and paid
$1.2 billion in dividends, returning over $4 billion to shareholders.
CEO Safra Catz emphasised that Oracle's investment in OCI AI infrastructure,
including over 40,000 NVIDIA H200 GPUs and partnerships with Google, Meta, and xAI,
positions Oracle as the partner of choice for enterprise AI deployments in 2026 and beyond.
"""
# ── Build the summarisation request ──────────────────────────────────────────
messages = [
SystemMessage(content=(
"You are a financial analyst assistant. "
"Produce concise, factual summaries with bullet points. "
"Focus on numbers and key business outcomes."
)),
HumanMessage(content=(
f"Please summarise the following financial report into:\n"
f"1. Three key takeaways (bullet points)\n"
f"2. One sentence of overall assessment\n\n"
f"Document:\n{long_document}"
))
]
print("📋 SUMMARISING QUARTERLY EARNINGS REPORT...")
print("-" * 60)
response = summariser_llm.invoke(messages)
print(response.content)
Example output:
📋 SUMMARISING QUARTERLY EARNINGS REPORT... ------------------------------------------------------------ **Key Takeaways:** • **Revenue Growth:** Total Q3 FY2026 revenue hit $16.9B (+17% YoY), led by cloud infrastructure with OCI growing an exceptional 49% YoY to $2.7B — driven by surging enterprise demand for AI and GPU compute. • **Database Strength:** Database and cloud services rose 15% to $4.1B, with Autonomous Database surpassing 28,000 paying customers, reflecting strong adoption of Oracle Database 23ai and AI Vector Search capabilities. • **Shareholder Returns:** Net income of $3.4B ($1.22 EPS) with over $4B returned to shareholders through $3.1B in buybacks and $1.2B in dividends. **Overall Assessment:** Oracle delivered strong broad-based growth in Q3 FY2026, with OCI AI infrastructure emerging as the clear growth engine, supported by an extensive GPU partnership ecosystem positioning Oracle as a leading enterprise AI platform.
A comprehensive 6-paragraph report distilled into 3 precise bullet points and one clear assessment sentence. That's the power of SOTA NLP models applied to document intelligence! 🎯
Use Case 3: RAG — Ask Questions About YOUR Own Documents 🔍
What Is RAG?
RAG stands for Retrieval-Augmented Generation. This is the most important pattern in enterprise AI. Let me explain with a simple story:
Imagine you hire a brilliant consultant (the LLM). They are incredibly smart and know everything from their years of study. But they don't know anything about YOUR company's specific internal policies, products, or data — because that was never in their textbooks!
RAG solves this by giving the consultant a relevant excerpt from your company's documents BEFORE they answer. It has two steps:
- Retrieve → Search your private documents to find the most relevant passages for the user's question
- Generate → Give those passages to the LLM as context, so it can answer accurately using YOUR data
WITHOUT RAG:
User: "What is our company's parental leave policy?"
LLM: "Parental leave policies typically range from 8-16 weeks..."
(generic answer based on training data — may be wrong for your company!)
WITH RAG:
Step 1: Search your HR policy document for "parental leave"
Step 2: Find: "Section 4.2: All full-time employees receive 18 weeks paid parental leave..."
Step 3: Give this to the LLM as context
LLM: "According to your HR policy (Section 4.2), full-time employees receive
18 weeks of paid parental leave..."
(accurate answer grounded in YOUR actual policy document!)
This is a complete RAG pipeline that does the following:
- Loads some company knowledge (product FAQ text)
- Converts that knowledge into mathematical vectors using Cohere Embed v4 (the embedding model)
- Stores those vectors in a FAISS vector database for fast searching
- When a user asks a question, searches the database for the most relevant knowledge
- Sends the found knowledge + user's question to the Llama 4 model to generate an accurate answer
from langchain_oci import ChatOCIGenAI, OCIGenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
OCI_ENDPOINT = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com"
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
# ══════════════════════════════════════════════════════════════════════════════
# PART A: CREATE THE KNOWLEDGE BASE FROM YOUR COMPANY'S DOCUMENTS
# ══════════════════════════════════════════════════════════════════════════════
# ── Step 1: Your company's knowledge (in production, load from PDF/DB/Wiki) ──
# This represents chunks of your company's internal documentation
company_knowledge = [
"Our Standard warranty covers all hardware defects for 24 months from purchase date. "
"Software bugs are covered for 12 months. Accidental damage is NOT covered under Standard warranty.",
"Premium Support customers receive 24/7 phone support with guaranteed 4-hour response time. "
"Standard Support customers receive email support with 48-hour response time on business days only.",
"Refund Policy: Full refund within 30 days of purchase if the product is returned in original packaging. "
"After 30 days, store credit only. After 90 days, no refunds are available.",
"To reset your password: Go to Settings → Security → Change Password. "
"You will receive an OTP on your registered email. "
"If you do not receive the OTP within 5 minutes, check your spam folder.",
"Our product supports Windows 10 and above, macOS 12 and above, and Ubuntu 20.04 LTS. "
"Linux Mint and Fedora are not officially supported but may work.",
"Bulk orders of 50+ units qualify for a 15% discount. "
"Orders above 200 units qualify for 25% discount plus free shipping. "
"Contact sales@company.com to place bulk orders.",
]
# ── Step 2: Create vector embeddings using Cohere Embed v4 ────────────────────
# Cohere Embed v4 converts each knowledge chunk into a list of numbers (a vector)
# that captures the MEANING of that text mathematically.
# Similar meaning = similar numbers. This enables semantic search!
print("🔢 Step 1: Converting knowledge into vectors (embeddings)...")
embedding_model = OCIGenAIEmbeddings(
model_id="cohere.embed-v4.0", # Latest Cohere multimodal embedding model
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={"input_type": "search_document"} # Tell it these are documents to index
)
# ── Step 3: Store vectors in FAISS vector database ────────────────────────────
# FAISS is a super-fast search engine for vectors.
# It can search through millions of vectors in milliseconds!
vector_store = FAISS.from_texts(
texts=company_knowledge,
embedding=embedding_model
)
retriever = vector_store.as_retriever(
search_kwargs={"k": 2} # When searching, return the 2 most relevant chunks
)
print("✅ Knowledge base created! Ready to answer questions.\n")
# ══════════════════════════════════════════════════════════════════════════════
# PART B: BUILD THE RAG CHAIN — RETRIEVE + GENERATE
# ══════════════════════════════════════════════════════════════════════════════
# ── Step 4: Set up the language model ────────────────────────────────────────
rag_llm = ChatOCIGenAI(
model_id="meta.llama-4-scout-17b-16e-instruct", # Efficient Llama 4 model
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={"temperature": 0.1, "max_tokens": 400} # Low temp = precise answers
)
# ── Step 5: Build the prompt template for RAG ────────────────────────────────
# This template tells the LLM to answer ONLY from the provided context
rag_prompt = ChatPromptTemplate.from_template("""
You are a helpful customer support assistant.
Answer the user's question based ONLY on the information provided in the context below.
If the answer is not in the context, say: "I don't have information about that in our knowledge base."
Do not make up information.
Context from our knowledge base:
{context}
Customer's Question:
{question}
Answer:""")
# ── Step 6: Build the complete RAG chain ─────────────────────────────────────
# This wires everything together: question → retrieve → generate → answer
def format_docs(docs):
"""Join retrieved document chunks into a single context string."""
return "\n\n".join(doc.page_content for doc in docs)
rag_chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| rag_prompt
| rag_llm
| StrOutputParser()
)
# ══════════════════════════════════════════════════════════════════════════════
# PART C: TEST WITH CUSTOMER QUESTIONS
# ══════════════════════════════════════════════════════════════════════════════
test_questions = [
"What is covered under the standard warranty?",
"Can I get a refund after 60 days?",
"What discount do I get for ordering 100 units?",
"Does your product work on Fedora Linux?",
]
print(" ASKING QUESTIONS USING THE RAG CHATBOT:")
print("=" * 60)
for question in test_questions:
print(f"\n❓ Customer: {question}")
answer = rag_chain.invoke(question)
print(f" Support Bot: {answer}")
print("-" * 50)
Example output:
🔢 Step 1: Converting knowledge into vectors (embeddings)... ✅ Knowledge base created! Ready to answer questions. 🤖 ASKING QUESTIONS USING THE RAG CHATBOT: ============================================================ ❓ Customer: What is covered under the standard warranty? 🤖 Support Bot: The Standard warranty covers all hardware defects for 24 months from your purchase date. Software bugs are covered for 12 months. Please note that accidental damage is not covered under the Standard warranty. -------------------------------------------------- ❓ Customer: Can I get a refund after 60 days? 🤖 Support Bot: After 60 days, you are no longer eligible for a full refund. Our policy allows full refunds within 30 days of purchase (in original packaging). Between 30 and 90 days, store credit is available. After 90 days, no refunds are available in any form. -------------------------------------------------- ❓ Customer: What discount do I get for ordering 100 units? 🤖 Support Bot: For an order of 100 units, you would qualify for our 200+ unit tier — wait, 100 units falls in the 50+ category, so you'd receive a 15% discount. For 25% discount plus free shipping, you would need an order of 200+ units. To place a bulk order, contact sales@company.com. -------------------------------------------------- ❓ Customer: Does your product work on Fedora Linux? 🤖 Support Bot: Fedora is not an officially supported operating system. Our product officially supports Windows 10+, macOS 12+, and Ubuntu 20.04 LTS. Fedora may work, but we cannot provide official support for it. --------------------------------------------------
The chatbot answered every question accurately using ONLY the company knowledge provided — not by making things up! That's the power and safety of RAG. 🎯
Use Case 4: Semantic Search with Embeddings 🔢
What Are Embeddings?
An embedding is a way of converting text into a list of numbers (a vector) that captures the meaning of that text. Two sentences that mean the same thing will have very similar vectors — even if they use completely different words.
KEYWORD SEARCH (old way) vs SEMANTIC SEARCH (new way): Query: "car won't start" KEYWORD SEARCH finds: ✅ "my car won't start in the morning" (exact keywords match) ❌ "vehicle fails to ignite" (same meaning, different words — NOT found!) ❌ "automobile refuses to turn on" (same meaning, different words — NOT found!) SEMANTIC SEARCH with Embeddings finds: ✅ "my car won't start in the morning" (same meaning — found!) ✅ "vehicle fails to ignite" (same meaning — found!) ✅ "automobile refuses to turn on" (same meaning — found!) ✅ "engine not responding on cold days" (related meaning — also found!) Embeddings understand MEANING, not just keywords. Revolutionary!
This code converts 6 job description sentences into vectors using Cohere Embed v4. Then it converts a search query into a vector. Finally, it finds which job descriptions are "closest" to the query in meaning — even if the words are different. This is exactly how LinkedIn's "Jobs You Might Like" or Indeed's recommendation engine works!
from langchain_oci import OCIGenAIEmbeddings
import numpy as np
OCI_ENDPOINT = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com"
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
# ── Step 1: Initialise the embedding model ────────────────────────────────────
# Cohere Embed v4 converts text into 1024-dimensional vectors (lists of 1024 numbers)
# that capture the full meaning and context of the text.
embedding_model = OCIGenAIEmbeddings(
model_id="cohere.embed-v4.0",
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={"input_type": "search_document"}
)
# ── Step 2: Job descriptions to search through ───────────────────────────────
job_descriptions = [
"Senior Python Developer — Build scalable backend APIs using Django and FastAPI.
5+ years experience required.", "Machine Learning Engineer — Design and train
neural networks for NLP tasks. PyTorch and Hugging Face experience essential.",
"DevOps Engineer — Manage CI/CD pipelines, Kubernetes clusters, and cloud
infrastructure on AWS and OCI.", "Data Analyst — Analyse sales data,
build dashboards in Tableau, and deliver weekly business reports.",
"AI Research Scientist — Publish research on large language models,
transformers, and RLHF techniques.", "Frontend Developer — Build
responsive web UIs using React.js and TypeScript. UX design skills a plus.",
]
# ── Step 3: Create embeddings for all job descriptions ────────────────────────
print("🔢 Creating embeddings for job descriptions...")
doc_embeddings = embedding_model.embed_documents(job_descriptions)
print(f"✅ {len(doc_embeddings)} embeddings created. Each vector has {len(doc_embeddings[0])} dimensions.\n")
# ── Step 4: User's search query ──────────────────────────────────────────────
search_query = "I want to work on AI and deep learning research"
# Create embedding for the query (using "search_query" input type)
query_embedding_model = OCIGenAIEmbeddings(
model_id="cohere.embed-v4.0",
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={"input_type": "search_query"}
)
query_vector = query_embedding_model.embed_query(search_query)
# ── Step 5: Calculate cosine similarity ───────────────────────────────────────
# Cosine similarity measures how similar two vectors are.
# 1.0 = identical meaning, 0.0 = completely unrelated, -1.0 = opposite
def cosine_similarity(vec_a, vec_b):
"""Calculate similarity between two vectors. Returns a score from -1 to 1."""
vec_a = np.array(vec_a)
vec_b = np.array(vec_b)
return np.dot(vec_a, vec_b) / (np.linalg.norm(vec_a) * np.linalg.norm(vec_b))
# Score every job description against the query
scores = []
for i, doc_vec in enumerate(doc_embeddings):
similarity = cosine_similarity(query_vector, doc_vec)
scores.append((similarity, job_descriptions[i]))
# ── Step 6: Show results sorted by relevance ─────────────────────────────────
scores.sort(key=lambda x: x[0], reverse=True) # Highest similarity first
print(f"🔍 Search Query: \"{search_query}\"")
print("=" * 65)
print("📊 Job descriptions ranked by semantic similarity:\n")
for rank, (score, description) in enumerate(scores, start=1):
bar_length = int(score * 30)
bar = "▓" * bar_length + "░" * (30 - bar_length)
print(f" Rank {rank}: [{bar}] {score:.3f}")
print(f" {description[:70]}...")
print()
Example output:
🔢 Creating embeddings for job descriptions...
✅ 6 embeddings created. Each vector has 1024 dimensions.
🔍 Search Query: "I want to work on AI and deep learning research"
=================================================================
📊 Job descriptions ranked by semantic similarity:
Rank 1: [▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░] 0.812
AI Research Scientist — Publish research on large language models, transformers...
Rank 2: [▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░░] 0.741
Machine Learning Engineer — Design and train neural networks for NLP tasks. Py...
Rank 3: [▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░░░░░░░░░░░] 0.431
Senior Python Developer — Build scalable backend APIs using Django and FastAPI...
Rank 4: [▓▓▓▓▓▓▓▓▓░░░░░░░░░░░░░░░░░░░░░] 0.342
DevOps Engineer — Manage CI/CD pipelines, Kubernetes clusters, and cloud infr...
Rank 5: [▓▓▓▓▓▓▓░░░░░░░░░░░░░░░░░░░░░░░] 0.278
Data Analyst — Analyse sales data, build dashboards in Tableau, and deliver w...
Rank 6: [▓▓▓▓▓░░░░░░░░░░░░░░░░░░░░░░░░░] 0.218
Frontend Developer — Build responsive web UIs using React.js and TypeScript. ...
"AI Research Scientist" ranked #1 and "Machine Learning Engineer" ranked #2 — both directly relevant to "AI and deep learning research". "Frontend Developer" ranked last — correct! It has nothing to do with AI research. Semantic search perfectly understood the meaning of the query! 🏆
Use Case 5: Prompt Engineering — Getting the Best Results 🎨
The way you phrase your input (called a prompt) dramatically affects the quality of the output. Prompt Engineering is the art and science of writing effective prompts. It's like knowing the right words to say to get the best help from a very capable assistant.
This demonstrates three levels of prompt quality for the same task. We compare a "bad" vague prompt, a "good" structured prompt, and a "best" few-shot prompt (where you show the model examples of what you want before asking it to do the real task). The difference in output quality is dramatic!
from langchain_oci import ChatOCIGenAI
from langchain_core.messages import HumanMessage, SystemMessage
OCI_ENDPOINT = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com"
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
llm = ChatOCIGenAI(
model_id="cohere.command-a-03-2025",
service_endpoint=OCI_ENDPOINT,
compartment_id=COMPARTMENT_ID,
model_kwargs={"temperature": 0.3, "max_tokens": 300}
)
product_review = (
"I bought this Bluetooth speaker 3 weeks ago. Sound quality is amazing and "
"the bass is deep and punchy. Battery lasts all day easily. My only complaint "
"is the Bluetooth pairing takes about 15 seconds — a bit slow. But overall "
"I'm very happy with it and would definitely recommend to friends."
)
# ── LEVEL 1: Bad prompt ───────────────────────────────────────────────────────
print("❌ LEVEL 1 — VAGUE PROMPT:")
bad_prompt = [HumanMessage(content=f"Analyse this review:\n\n{product_review}")]
bad_response = llm.invoke(bad_prompt)
print(bad_response.content[:200] + "...\n")
# ── LEVEL 2: Good structured prompt ──────────────────────────────────────────
print("✅ LEVEL 2 — STRUCTURED PROMPT:")
good_prompt = [
SystemMessage(content="You are a product review analyst. Always respond in JSON format."),
HumanMessage(content=(
f"Analyse this product review. Extract:\n"
f"1. Overall sentiment: positive/negative/neutral\n"
f"2. Positives mentioned (list)\n"
f"3. Negatives mentioned (list)\n"
f"4. Would they recommend? yes/no\n"
f"5. Confidence score 0-100\n\n"
f"Review: {product_review}"
))
]
good_response = llm.invoke(good_prompt)
print(good_response.content + "\n")
# ── LEVEL 3: Best few-shot prompt (show examples first!) ─────────────────────
print("🏆 LEVEL 3 — FEW-SHOT PROMPT (With Examples):")
fewshot_prompt = [
SystemMessage(content="You are a product review analyst. Extract insights in the exact JSON format shown."),
HumanMessage(content="""
Here are two example analyses:
EXAMPLE 1:
Review: "Great laptop! Super fast and light. Screen is beautiful. A bit pricey though."
Output: {"sentiment": "positive", "positives": ["fast", "lightweight", "beautiful screen"],
"negatives": ["expensive"], "recommend": "yes", "confidence": 87}
EXAMPLE 2:
Review: "Terrible phone. Crashes constantly, battery dies in 3 hours. Camera is ok I guess."
Output: {"sentiment": "negative", "positives": ["decent camera"],
"negatives": ["frequent crashes", "poor battery life"], "recommend": "no", "confidence": 92}
Now analyse this review in EXACTLY the same JSON format:
Review: """ + product_review)
]
best_response = llm.invoke(fewshot_prompt)
print(best_response.content)
Example output:
❌ LEVEL 1 — VAGUE PROMPT:
The review discusses a Bluetooth speaker with positive overall tone. The customer
mentions good sound quality and battery life but notes slow Bluetooth pairing...
✅ LEVEL 2 — STRUCTURED PROMPT:
{
"overall_sentiment": "positive",
"positives": ["amazing sound quality", "deep bass", "all-day battery life"],
"negatives": ["slow Bluetooth pairing (15 seconds)"],
"would_recommend": "yes",
"confidence_score": 88
}
🏆 LEVEL 3 — FEW-SHOT PROMPT (With Examples):
{"sentiment": "positive", "positives": ["amazing sound quality", "deep punchy bass",
"all-day battery life"], "negatives": ["slow Bluetooth pairing (15 seconds)"],
"recommend": "yes", "confidence": 91}
Level 1 produced loose prose. Level 2 produced structured JSON. Level 3 produced perfectly formatted JSON that exactly matches the required schema — ready to be parsed by code and stored in a database automatically! 🎯
Fine-Tuning — Customising SOTA Models on OCI 🔧
Pre-trained models are incredibly capable out of the box. But for highly specialised domains (medical, legal, finance, manufacturing), you can fine-tune them on YOUR specific data to make them even better.
FINE-TUNING vs RAG — When to Use Which? RAG (Retrieval-Augmented Generation): ✅ Use when: Your data changes frequently (new policies, new products) ✅ Use when: You need the model to cite sources ✅ Use when: You have a large, varied knowledge base ✅ Easier to set up — no training required ❌ Slightly slower (must retrieve first, then generate) Fine-Tuning: ✅ Use when: You need the model to BEHAVE differently (not just know more) ✅ Use when: You have a specific output format or style required consistently ✅ Use when: The task is highly specialised (medical coding, legal clauses) ✅ Faster inference — the model already knows the format ❌ Requires labelled training data (prompt + ideal completion pairs) ❌ Data doesn't update automatically (must retrain for new data)
OCI supports fine-tuning on two model families:
- 🔵 Cohere Command R — Fine-tuned using T-Few (fast, efficient) or Vanilla (full fine-tuning with granular control over layers)
- 🦙 Meta Llama 3.3 70B Instruct — Fine-tuned using LoRA (Low-Rank Adaptation): only trains small "adapter" matrices rather than all 70 billion parameters — making it 10x more efficient!
Your training data must be a JSONL file (JSON Lines — one JSON object per line).
Each line must have exactly two keys:
prompt (the input) and completion (the ideal output).Example line:
{"prompt": "Classify this support ticket: 'My app crashes on startup'", "completion": "Category: Technical Bug | Priority: High"}Minimum recommended: 50 training examples. For best results: 200+ examples.
Real-World Industry Use Cases — SOTA NLP on OCI
Here are real patterns deployed by enterprises :
-
🏦 Banking — Intelligent Document Processing
A bank receives 10,000 loan applications per month as unstructured PDF documents. OCI Cohere Command A reads each document, extracts applicant details, income figures, and risk factors, and structures them into a database entry — automatically. Processing time drops from 3 days to 4 minutes per application. -
🏥 Healthcare — Clinical Note Summarisation
Doctors dictate notes after every patient visit. A fine-tuned Llama 3.3 70B model reads each note and generates a structured SOAP summary (Subjective, Objective, Assessment, Plan) for the Electronic Health Record. Saves each doctor 45 minutes of admin work per day. -
⚖️ Legal — Contract Intelligence
A law firm deploys RAG over 50 years of contracts. Lawyers ask natural language questions: "Show me all contracts where the indemnification clause exceeds $5 million and expires before June 2027." The system retrieves and analyses relevant contracts in seconds — work that would take a paralegal 2 weeks. -
🛒 E-Commerce — Multilingual Customer Support
A global retailer serves customers in 28 countries. Cohere Command A's multilingual capabilities allow a single RAG chatbot to answer questions in the customer's native language — pulling from the same English knowledge base and responding fluently in French, German, Spanish, Japanese, or Arabic. -
💻 Software Development — AI Code Assistant
An enterprise developer team integrates Llama 4 Maverick via OCI API into their IDE. Developers write comments in plain English; the AI generates production-ready Java, Python, or PL/SQL code. Code review time reduces by 40% and junior developer productivity doubles.
Quick Summary 📝
What we learned about SOTA NLP Models on OCI:
- SOTA NLP Models → The world's most powerful AI language models, available through a single OCI API endpoint
- OCI Model Catalogue→ Cohere Command A/R+/R, Meta Llama 4 Maverick/Scout/3.3, Google Gemini 2.5, xAI Grok, OpenAI GPT — all on one platform
- LLM Parameters → Billions of learned settings. More = smarter (but costlier). Temperature, top-p, and max_tokens control response style and length
- Chat / Text Generation → Natural conversation and content creation — the most fundamental NLP task
- Summarisation → Condense long documents to key points — use low temperature (0.1) for precision
- RAG → The killer enterprise pattern: ground LLM answers in YOUR private documents using retrieval + generation
- Embeddings → Convert text to vectors that capture meaning. Power semantic search, recommendations, and RAG retrieval
- Prompt Engineering → How you phrase your prompt massively affects output quality — structured prompts and few-shot examples dramatically improve results
- Fine-Tuning → Customise Cohere or Llama models on your specific training data for domain-specific performance gains
Happy building! 🧠☁️
Comments
Post a Comment