A GenAI Pipeline is like a factory assembly line — raw data goes in at one end, and intelligent AI-powered answers come out the other end. Every step in between is carefully designed, just like a real factory! 🏭
What Even IS a GenAI Pipeline?
Imagine you want to build a smart assistant for a hospital. 🏥 Doctors can ask it questions like: "What is the latest treatment for diabetes?" and it gives a perfect, accurate answer.
But how does it actually work behind the scenes? A magic button? No! It is a carefully designed pipeline — a series of connected steps that work together like a relay race. 🏃♂️🏃♀️🏃
Each runner passes the baton to the next. If one runner drops it — the whole race is lost. That is why understanding every step of the pipeline matters.
THE BIG PICTURE — End-to-End GenAI Pipeline on OCI ┌─────────────────────────────────────────────────────────────────────────────────┐ │ │ │ 📦 RAW DATA 🔧 PROCESS 🧠 UNDERSTAND 💬 ANSWER │ │ (PDFs, DBs, ───► (Clean & ───► (Embeddings ───► (LLM gives │ │ Images, Chunk it) + Vector DB) the reply) │ │ APIs, Logs) │ │ │ │ Step 1 Step 2 Step 3 & 4 Step 5 & 6 │ │ Ingestion Preprocessing RAG + Vector Store LLM Generation │ │ │ └─────────────────────────────────────────────────────────────────────────────────┘ All of this runs on OCI (Oracle Cloud Infrastructure) ☁️
A GenAI Pipeline = A sequence of automated steps that takes raw data, processes it, stores it intelligently, and uses a Large Language Model (LLM) to generate smart, context-aware answers.
🗺️ The Full Journey — All 7 Steps
Think of building a GenAI Pipeline like building a water purification plant. 💧 Dirty river water (raw data) comes in. Clean, safe drinking water (AI answers) goes out. There are several purification stages in between — each one important!
- 🪣 Step 1: Data Ingestion — Collect the raw water (data)
- 🧹 Step 2: Preprocessing — Filter and clean it
- 🧪 Step 3: Embeddings — Understand the meaning of every drop
- 🗄️ Step 4: Vector Database — Store in smart memory tanks
- 🔍 Step 5: Query & Retrieval — Find the right water when asked
- 🤖 Step 6: LLM Generation — Pour out the perfect clean answer
- 🖼️ Step 7: Multi-modal — Handle text, images, AND tables together
🪣 Step 1 — Data Ingestion: Collecting Your Raw Data
Before any AI can be smart, it needs lots and lots of data. But where does that data come from?
In a real company, data lives everywhere — in PDF files, spreadsheets, databases, websites, emails, images, and even audio recordings. Data Ingestion is the step where we collect all of this from different places and bring it into one central location.
Think of Data Ingestion like a supermarket delivery truck. 🚚 It visits farms, factories, and warehouses — and brings all the different food items to the supermarket in one trip. The supermarket is your OCI Object Storage.
WHERE DOES DATA COME FROM?
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 📄 PDF/Word │ │ 🗃️ Databases│ │ 🌐 Web APIs │ │ 🖼️ Images │
│ Documents │ │ (Oracle DB, │ │ (REST JSON) │ │ Videos, │
│ Reports │ │ MySQL) │ │ │ │ Audio │
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │ │
└──────────────────┴───────────────────┴───────────────────┘
│
▼
┌─────────────────────────┐
│ ☁️ OCI Object Storage │
│ (Your Central Data │
│ Landing Zone) │
└─────────────────────────┘
🛠️ OCI Tools Used for Data Ingestion
- ☁️ OCI Object Storage — The main bucket to store all raw files (PDFs, CSVs, images)
- 🔄 OCI Data Integration — A drag-and-drop tool to pull data from many sources
- ⚡ OCI Functions — Small automated programs that watch for new files and trigger the pipeline
- 🌐 OCI API Gateway — Collects data from external REST APIs automatically
Never dump all your data in one giant messy folder with no structure. Always organise data in separate OCI buckets by type: one for PDFs, one for images, one for database exports. A messy kitchen makes terrible food! 🍳
Use a "Bronze / Silver / Gold" bucket pattern on OCI Object Storage:
- 🟤 Bronze Bucket — Raw, untouched data (exactly as it arrived)
- ⚪ Silver Bucket — Cleaned and validated data
- 🟡 Gold Bucket — Final, ready-to-use data for the AI
Now let us look at a real code example. Before that — here is what the code below actually does:
It connects to your OCI account, opens a storage bucket, and uploads a file from your computer to the cloud. Just like dragging a file into Google Drive — but done automatically by a program! 📁➡️☁️
# ---------------------------------------------------------------
# PURPOSE: Upload a local file (e.g., hospital_report.pdf)
# to OCI Object Storage Bronze Bucket automatically.
# ---------------------------------------------------------------
import oci
# Step 1: Load your OCI credentials from the config file
config = oci.config.from_file("~/.oci/config")
# Step 2: Create a connection to the Object Storage service
object_storage = oci.object_storage.ObjectStorageClient(config)
# Step 3: Define where to upload (your bucket name and file)
namespace = object_storage.get_namespace().data
bucket_name = "bronze-raw-data"
file_name = "hospital_report.pdf"
# Step 4: Read the file from your local computer
with open(file_name, "rb") as f:
file_content = f.read()
# Step 5: Upload it to OCI Object Storage
object_storage.put_object(
namespace_name = namespace,
bucket_name = bucket_name,
object_name = file_name,
put_object_body= file_content
)
print(f"✅ File '{file_name}' uploaded to OCI bucket '{bucket_name}' successfully!")
🧹 Step 2 — Preprocessing: Cleaning and Preparing the Data
Raw data is almost never clean. It has typos, missing values, strange characters, and messy formatting. If you feed dirty data to your AI — it will give you dirty answers. 🗑️
Preprocessing is like washing vegetables before cooking them. 🥦 No matter how good your recipe is, if the vegetables are dirty — the dish will taste bad.
WHAT PREPROCESSING DOES:
RAW TEXT (messy): CLEAN TEXT (after preprocessing):
┌──────────────────────────┐ ┌──────────────────────────────────┐
│ "Patient naame: John!!! │ │ "Patient name: John. │
│ Diagnosiss: diABEtes. │ ──► │ Diagnosis: Diabetes. │
│ Age: 45 yrs old ???? │ │ Age: 45." │
│ [PAGE BREAK][HEADER]... │ │ │
└──────────────────────────┘ └──────────────────────────────────┘
Garbage In ❌ Quality Out ✅
🔧 What Happens During Preprocessing?
- 📄 Text Extraction — Pull text out of PDFs, Word docs, HTML pages
- 🧽 Noise Removal — Remove headers, footers, page numbers, special characters
- ✂️ Chunking — Break long documents into small, manageable pieces (chunks)
- 🌍 Language Detection — Identify if text is English, Hindi, Spanish, etc.
- 🔗 Deduplication — Remove duplicate paragraphs or repeated content
An AI cannot read a 500-page book in one go — just like you cannot eat an entire pizza in one bite! 🍕 Chunking breaks the document into slices (usually 200–500 words per chunk). Each chunk is processed and stored separately. This makes retrieval much faster and more accurate.
It reads a PDF file, extracts all the text from it, and then cuts that text into small overlapping pieces called "chunks." These chunks are what the AI will actually read and understand! 📄✂️
# ---------------------------------------------------------------
# PURPOSE: Extract text from a PDF and split it into
# small chunks (like cutting a book into pages).
# ---------------------------------------------------------------
from PyPDF2 import PdfReader
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Step 1: Open and read the PDF file
reader = PdfReader("hospital_report.pdf")
raw_text = ""
for page in reader.pages:
raw_text += page.extract_text()
print(f"📄 Total characters extracted: {len(raw_text)}")
# Step 2: Set up the text splitter
# chunk_size = how many characters per chunk (like slices of bread 🍞)
# chunk_overlap = how many characters overlap between chunks
# (so we don't lose meaning at the edges)
splitter = RecursiveCharacterTextSplitter(
chunk_size = 500, # Each chunk = ~500 characters
chunk_overlap = 50 # 50 characters shared with next chunk
)
# Step 3: Split the text into chunks
chunks = splitter.split_text(raw_text)
print(f"✅ Total chunks created: {len(chunks)}")
print(f"\n📌 First chunk preview:\n{chunks[0]}")
🧪 Step 3 — Embeddings: Teaching AI to "Understand" Words
Here is the most magical step of the entire pipeline. 🪄 Computers don't actually understand words like you and I do. They only understand numbers.
So how do we make a computer understand that "car" and "automobile" mean the same thing? Or that "happy" is the opposite of "sad"? The answer is Embeddings.
An embedding is a list of numbers that captures the meaning of a word or sentence. Words with similar meanings get similar numbers. Words with different meanings get very different numbers.
Think of it like a GPS coordinate for meaning. 🗺️ "King" and "Queen" have GPS coordinates that are very close to each other. "King" and "Pizza" have GPS coordinates far apart. The AI navigates using these coordinates!
HOW EMBEDDINGS WORK: Word/Sentence Embedding (list of numbers) ┌───────────────────┐ ┌──────────────────────────────────────────┐ │ "The patient has │ ─────► │ [0.23, -0.81, 0.45, 0.12, -0.33, ...] │ │ high blood │ │ (1536 numbers total — each capturing │ │ pressure." │ │ a different aspect of meaning) │ └───────────────────┘ └──────────────────────────────────────────┘ "high blood pressure" ──► [0.23, -0.81, 0.45 ...] "hypertension" ──► [0.22, -0.80, 0.44 ...] ← VERY SIMILAR! ✅ "chocolate cake" ──► [0.91, 0.15, -0.72 ...] ← VERY DIFFERENT! ❌
🔌 OCI Generative AI — Embedding Models Available
- 🟢 Cohere Embed v3 — Best for multilingual text (supports 100+ languages)
- 🔵 Cohere Embed English v3 — Fastest for English-only content
- 🟣 Custom Fine-tuned Models — You can train your own embedding model on OCI
It sends a text chunk to the OCI Generative AI service and gets back a list of numbers (the embedding) that represents the meaning of that text. Think of it like sending a letter and getting back the GPS coordinates of its meaning! 📬🗺️
# ---------------------------------------------------------------
# PURPOSE: Convert text chunks into embeddings (lists of numbers)
# using OCI Generative AI Cohere Embed model.
# ---------------------------------------------------------------
import oci
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import (
EmbedTextDetails, OnDemandServingMode
)
# Step 1: Connect to OCI Generative AI service
config = oci.config.from_file("~/.oci/config")
client = GenerativeAiInferenceClient(config)
# Step 2: The text chunk we want to convert to embedding
text_chunk = "The patient has high blood pressure and was prescribed lisinopril."
# Step 3: Create the embedding request
embed_request = EmbedTextDetails(
inputs = [text_chunk], # Our text chunk
serving_mode = OnDemandServingMode(
model_id = "cohere.embed-multilingual-v3" # OCI Embedding Model
),
compartment_id = "ocid1.compartment.oc1..xxxxx", # Your OCI compartment
input_type = "SEARCH_DOCUMENT" # We are embedding a document
)
# Step 4: Send to OCI and get the embedding back
response = client.embed_text(embed_request)
embedding = response.data.embeddings[0]
print(f"✅ Embedding created!")
print(f"📏 Number of dimensions: {len(embedding)}")
print(f"🔢 First 5 values: {embedding[:5]}")
# Output: [0.023, -0.415, 0.812, -0.104, 0.337]
🗄️ Step 4 — Vector Database: The AI's Smart Memory
Now that we have turned all our chunks into embeddings (lists of numbers), we need to store them somewhere special.
A regular database stores data in rows and columns (like a spreadsheet). But embeddings need a Vector Database — a special type of database that can store lists of numbers AND find similar ones at lightning speed. ⚡
Imagine a library with millions of books. 📚 A normal librarian (regular database) can only find a book if you give the exact title. A vector database librarian can find books that are similar in topic even if you don't know the exact title! You say: "I want something about space travel and adventure" — and they bring you 5 perfect matching books instantly. That is the power of vector search! 🚀
REGULAR DATABASE vs VECTOR DATABASE: Regular DB (SQL): Vector Database: ┌──────┬──────────────────────┐ ┌──────────────────────────────────────────┐ │ ID │ Text │ │ ID │ Text │ Embedding │ ├──────┼──────────────────────┤ ├─────┼─────────────────┼──────────────────┤ │ 1 │ "Diabetes info" │ │ 1 │ "Diabetes info"│ [0.2, -0.8 ...] │ │ 2 │ "Hypertension info" │ │ 2 │ "Hypertension" │ [0.3, -0.7 ...] │ └──────┴──────────────────────┘ └─────┴─────────────────┴──────────────────┘ Search: "blood sugar problem" Search: "blood sugar problem" Result: ❌ No exact match found! Result: ✅ Finds "Diabetes info" (similar meaning!)
🟡 Oracle Database 23ai — The Best Vector DB on OCI
The most powerful choice for a Vector Database on OCI is Oracle Database 23ai with built-in AI Vector Search.
Why is this special? Because Oracle 23ai is not just a vector database. It is a full relational database AND a vector database in one place. You can store your regular data AND your embeddings in the same database — no extra tools needed! 🎯
- 🔍 VECTOR data type — Store embeddings directly in Oracle tables
- ⚡ HNSW Index — Super-fast similarity search (finds similar chunks in milliseconds)
- 🔒 Enterprise Security — Built-in encryption, access control
- ☁️ Fully Managed on OCI — No server setup needed (Autonomous Database)
It creates a table in Oracle Database 23ai that can store text AND its embedding (numbers). Then it saves one chunk of text with its embedding into the table. Think of it like adding a book to the smart library along with its GPS location! 📚🗺️
# ---------------------------------------------------------------
# PURPOSE: Store text chunks AND their embeddings into
# Oracle Database 23ai (the smart vector library).
# ---------------------------------------------------------------
import oracledb
import array
# Step 1: Connect to Oracle Autonomous Database (OCI)
connection = oracledb.connect(
user = "ADMIN",
password = "YourSecurePassword123!",
dsn = "your_adb_high" # Connection string from OCI wallet
)
cursor = connection.cursor()
# Step 2: Create a table with VECTOR column (only needed once)
cursor.execute("""
CREATE TABLE IF NOT EXISTS hospital_knowledge (
id NUMBER GENERATED ALWAYS AS IDENTITY,
chunk_text CLOB,
embedding VECTOR(1536, FLOAT32), -- stores 1536 numbers per chunk
source_file VARCHAR2(200),
created_at TIMESTAMP DEFAULT SYSTIMESTAMP
)
""")
# Step 3: Our chunk text and its embedding (from Step 3)
chunk_text = "The patient has high blood pressure and was prescribed lisinopril."
embedding = [0.023, -0.415, 0.812, -0.104, 0.337] # (simplified — real ones have 1536 values)
# Step 4: Convert embedding to Oracle VECTOR format
embedding_array = array.array("f", embedding)
# Step 5: Insert into the table
cursor.execute("""
INSERT INTO hospital_knowledge (chunk_text, embedding, source_file)
VALUES (:1, :2, :3)
""", [chunk_text, embedding_array, "hospital_report.pdf"])
connection.commit()
print("✅ Chunk and embedding stored successfully in Oracle 23ai!")
cursor.close()
connection.close()
🔍 Step 5 — RAG Architecture and Query Retrieval Workflow
This is where everything comes together! 🎉 RAG stands for Retrieval-Augmented Generation.
When a user asks a question, the pipeline does two things before answering:
- 🔍 RETRIEVE — Find the most relevant chunks from the vector database
- ✍️ GENERATE — Give those chunks to the LLM and ask it to write an answer
A regular LLM only knows what it was trained on (data up to a certain date). It does NOT know your private company data, your hospital records, or your latest reports.
RAG injects your private data into the conversation dynamically. The LLM then uses that fresh, specific information to answer — no retraining needed! 🎯
THE COMPLETE RAG WORKFLOW:
USER ASKS: "What medication was prescribed for high blood pressure?"
│
▼
┌─────────────────────────────────────────────────────┐
│ STEP A: Convert the question into an embedding │
│ "What medication for high blood pressure?" │
│ ──► [0.025, -0.410, 0.805, ...] │
└────────────────────┬────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ STEP B: Search Vector DB for similar chunks │
│ Compare question embedding vs all stored embeddings│
│ Find TOP 3 most similar chunks ──► │
│ Chunk 1: "...prescribed lisinopril..." │
│ Chunk 2: "...blood pressure treatment..." │
│ Chunk 3: "...antihypertensive drugs..." │
└────────────────────┬────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ STEP C: Send question + chunks to LLM │
│ "Here is context: [chunk1][chunk2][chunk3] │
│ Now answer: What medication for blood pressure?" │
└────────────────────┬────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ LLM ANSWER: "Based on the patient records, │
│ lisinopril was prescribed for high blood pressure."│
└─────────────────────────────────────────────────────┘
When a user asks a question, this code converts the question into an embedding, searches the Oracle 23ai database for the 3 most similar chunks, and returns them ready to send to the LLM. It is the "search" step! 🔍
# ---------------------------------------------------------------
# PURPOSE: Take a user question, convert it to an embedding,
# and find the TOP 3 most relevant chunks from Oracle 23ai.
# This is the "Retrieval" part of RAG.
# ---------------------------------------------------------------
import oracledb
import array
import oci
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import EmbedTextDetails, OnDemandServingMode
def get_embedding(text: str) -> list:
"""Convert any text into an embedding using OCI."""
config = oci.config.from_file("~/.oci/config")
client = GenerativeAiInferenceClient(config)
req = EmbedTextDetails(
inputs = [text],
serving_mode = OnDemandServingMode(model_id="cohere.embed-multilingual-v3"),
compartment_id = "ocid1.compartment.oc1..xxxxx",
input_type = "SEARCH_QUERY" # This time it is a query, not a document
)
response = client.embed_text(req)
return response.data.embeddings[0]
def retrieve_relevant_chunks(user_question: str, top_k: int = 3) -> list:
"""Find the top_k most relevant chunks for the user's question."""
# Step 1: Convert user question to embedding
question_embedding = get_embedding(user_question)
embedding_array = array.array("f", question_embedding)
# Step 2: Search Oracle 23ai using VECTOR_DISTANCE function
# VECTOR_DISTANCE = how far apart two embeddings are
# ORDER BY distance ASC = closest (most similar) first
connection = oracledb.connect(
user = "ADMIN",
password = "YourSecurePassword123!",
dsn = "your_adb_high"
)
cursor = connection.cursor()
cursor.execute("""
SELECT chunk_text,
VECTOR_DISTANCE(embedding, :query_vec, COSINE) AS similarity_score
FROM hospital_knowledge
ORDER BY similarity_score ASC -- closest match first
FETCH FIRST :top_k ROWS ONLY
""", query_vec=embedding_array, top_k=top_k)
results = [row[0] for row in cursor.fetchall()]
cursor.close()
connection.close()
return results
# --- TEST IT ---
question = "What medication was prescribed for high blood pressure?"
top_chunks = retrieve_relevant_chunks(question, top_k=3)
print(f"🔍 Question: {question}")
print(f"\n📌 Top {len(top_chunks)} relevant chunks found:")
for i, chunk in enumerate(top_chunks, 1):
print(f"\n Chunk {i}: {chunk[:150]}...")
🤖 Step 6 — Integrating LLMs for Generative Tasks
Now we have the relevant chunks retrieved from our vector database. The final step is to send these chunks — along with the user's question — to a Large Language Model (LLM) and let it write a perfect answer.
Think of the LLM as a very brilliant friend who reads extremely fast. 🧠 You hand them 3 pages of notes (the retrieved chunks) and ask a question. They read the notes, understand them, and write you a perfect, clear answer.
🤖 LLMs Available on OCI Generative AI Service
- 🦙 Meta Llama 3.1 / 3.2 / 3.3 — Powerful open-source models, great for enterprise
- 🟢 Cohere Command R+ — Optimized for RAG, very good at using context
- ⚡ Cohere Command R — Faster and lighter for simpler tasks
- 🔧 Custom Fine-tuned Models — Train your own model on OCI with your data
For RAG pipelines, use Cohere Command R+ on OCI — it is specifically designed to read context documents and give grounded, accurate answers. It is less likely to "hallucinate" (make things up) compared to other models. 🎯
It takes the user's question and the top matching chunks, builds a prompt (a carefully written instruction for the AI), sends it to OCI's LLM, and prints the final AI-generated answer. This is the "Generation" part of RAG! ✍️🤖
# ---------------------------------------------------------------
# PURPOSE: Send the user question + retrieved chunks to
# OCI Generative AI LLM and get a smart answer back.
# This is the "Generation" part of RAG.
# ---------------------------------------------------------------
import oci
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import (
ChatDetails, OnDemandServingMode,
CohereChatRequest, CohereMessage
)
def generate_answer(user_question: str, context_chunks: list) -> str:
"""Use OCI LLM to generate an answer based on retrieved context."""
config = oci.config.from_file("~/.oci/config")
client = GenerativeAiInferenceClient(config)
# Step 1: Combine all chunks into one "context" block
context = "\n\n---\n\n".join(context_chunks)
# Step 2: Build the prompt
# We tell the LLM exactly what role it plays and what to do
system_message = """You are a helpful medical assistant.
Answer the user's question using ONLY the context provided below.
If the answer is not in the context, say "I don't have that information."
Always be accurate and concise."""
user_message = f"""
CONTEXT (from hospital records):
{context}
QUESTION:
{user_question}
Please provide a clear and accurate answer based on the context above.
"""
# Step 3: Build the chat request for Cohere Command R+
chat_request = CohereChatRequest(
message = user_message,
preamble = system_message,
max_tokens = 500, # Maximum words in the reply
temperature = 0.1, # Low = more factual, High = more creative
is_stream = False # Get full answer at once
)
# Step 4: Send to OCI and get the response
response = client.chat(ChatDetails(
serving_mode = OnDemandServingMode(
model_id = "cohere.command-r-plus" # Our chosen LLM
),
compartment_id = "ocid1.compartment.oc1..xxxxx",
chat_request = chat_request
))
return response.data.chat_response.text
# --- PUTTING IT ALL TOGETHER ---
user_question = "What medication was prescribed for high blood pressure?"
context_chunks = retrieve_relevant_chunks(user_question, top_k=3) # From Step 5
final_answer = generate_answer(user_question, context_chunks)
print(f"❓ Question: {user_question}")
print(f"\n🤖 AI Answer:\n{final_answer}")
# Output:
# 🤖 AI Answer:
# "Based on the patient records, lisinopril was prescribed
# for the patient with high blood pressure."
🖼️ Step 7 — Multi-Modal Workflows: Text + Images + Structured Data
So far we have only worked with text. But the real world has many types of data — X-ray images, Excel spreadsheets, charts, scanned forms, and even audio recordings.
A Multi-modal Pipeline can understand and work with multiple types of data at the same time. It is like having a team where one person reads documents, another looks at images, and a third reads spreadsheets — and they all talk to each other to give you the best answer. 🤝
MULTI-MODAL PIPELINE FLOW:
Input Types:
┌────────────┐ ┌────────────┐ ┌────────────┐
│ 📄 Text │ │ 🖼️ Images │ │ 📊 Tables │
│ (PDFs, │ │ (X-rays, │ │ (Excel, │
│ Emails) │ │ Charts) │ │ CSV) │
└─────┬──────┘ └─────┬──────┘ └─────┬──────┘
│ │ │
▼ ▼ ▼
Text Image Captioning Table Parsing
Embeddings (Vision LLM) (Extract rows/cols)
│ │ │
└────────────────┴─────────────────┘
│
▼
┌──────────────────────┐
│ Unified Vector DB │
│ (Oracle 23ai) │
│ All embeddings │
│ stored together │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Multi-Modal LLM │
│ Gives one unified │
│ answer using ALL │
│ data types │
└──────────────────────┘
🖼️ How Does the AI "See" Images?
OCI Generative AI supports Vision-enabled LLMs like Llama 3.2 Vision. You can send an image directly to the model and ask questions about it!
It sends an X-ray image to the OCI Vision LLM and asks it to describe what it sees. The model "looks" at the image just like a doctor would, and writes a text description of what it found! 🏥🖼️➡️📝
# ---------------------------------------------------------------
# PURPOSE: Send a medical image (X-ray) to OCI's Vision LLM
# (Llama 3.2 Vision) and get a text description.
# Think of it as asking an AI doctor to "read" the image.
# ---------------------------------------------------------------
import oci
import base64
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import (
ChatDetails, OnDemandServingMode,
GenericChatRequest, UserMessage,
TextContent, ImageContent, ImageUrl
)
def analyze_medical_image(image_path: str, question: str) -> str:
"""Ask the Vision LLM to analyze an image and answer a question."""
config = oci.config.from_file("~/.oci/config")
client = GenerativeAiInferenceClient(config)
# Step 1: Read image and convert to base64
# (base64 is a way to send binary files as text over the internet)
with open(image_path, "rb") as img_file:
image_data = base64.b64encode(img_file.read()).decode("utf-8")
image_base64 = f"data:image/jpeg;base64,{image_data}"
# Step 2: Build a message with both the image AND the question
message = UserMessage(
content = [
ImageContent(
image_url = ImageUrl(url=image_base64) # The image
),
TextContent(
text = question # The question about the image
)
]
)
# Step 3: Create the chat request using Llama 3.2 Vision
chat_request = GenericChatRequest(
messages = [message],
max_tokens = 300,
temperature = 0.2
)
# Step 4: Send to OCI and get the visual analysis back
response = client.chat(ChatDetails(
serving_mode = OnDemandServingMode(
model_id = "meta.llama-3.2-90b-vision-instruct" # Vision LLM on OCI
),
compartment_id = "ocid1.compartment.oc1..xxxxx",
chat_request = chat_request
))
return response.data.chat_response.choices[0].message.content[0].text
# --- TEST IT ---
result = analyze_medical_image(
image_path = "chest_xray.jpg",
question = "Describe any abnormalities visible in this chest X-ray."
)
print(f"🖼️ Image Analysis Result:\n{result}")
# Output:
# 🖼️ Image Analysis Result:
# "The chest X-ray shows mild cardiomegaly (enlarged heart).
# Lung fields appear clear with no visible consolidation or effusion."
🏗️ The Complete End-to-End OCI Architecture
Now let us zoom out and see all 7 steps together on OCI. This is the full picture of a production-grade GenAI Pipeline that a real company would use.
COMPLETE END-TO-END GENAI PIPELINE ON OCI ┌────────────────────────────────────────────────────────────────────────────┐ │ │ │ DATA SOURCES │ │ 📄 PDFs 🗃️ Databases 🌐 APIs 🖼️ Images 📊 Spreadsheets │ │ │ │ │ ▼ │ │ ┌────────────────┐ INGESTION │ │ │ OCI Object │ ◄── OCI Data Integration + OCI Functions │ │ │ Storage │ (Bronze → Silver → Gold Buckets) │ │ │ (Bronze Bucket)│ │ │ └───────┬────────┘ │ │ │ │ │ ▼ │ │ ┌────────────────┐ PREPROCESSING │ │ │ OCI Data Flow │ ──► Text extract, clean, chunk │ │ │ (Apache Spark) │ (LangChain + OCI Compute) │ │ └───────┬────────┘ │ │ │ │ │ ▼ │ │ ┌────────────────┐ EMBEDDINGS │ │ │ OCI GenAI │ ──► Cohere Embed v3 Multilingual │ │ │ Embed Model │ Converts chunks → vectors │ │ └───────┬────────┘ │ │ │ │ │ ▼ │ │ ┌────────────────┐ VECTOR STORE │ │ │ Oracle DB 23ai │ ──► Stores text + VECTOR embeddings │ │ │ AI Vector │ HNSW Index for fast similarity search │ │ │ Search │ │ │ └───────┬────────┘ │ │ │ ◄────────── USER ASKS A QUESTION │ │ ▼ │ │ ┌────────────────┐ RETRIEVAL (RAG) │ │ │ OCI Functions │ ──► Question → embedding → similarity search │ │ │ (Query Engine) │ TOP 3 matching chunks retrieved │ │ └───────┬────────┘ │ │ │ │ │ ▼ │ │ ┌────────────────┐ GENERATION │ │ │ OCI GenAI │ ──► Cohere Command R+ or Llama 3.3 │ │ │ LLM Service │ Reads context + generates the answer │ │ └───────┬────────┘ │ │ │ │ │ ▼ │ │ 💬 FINAL ANSWER delivered to the user via OCI API Gateway │ │ │ └────────────────────────────────────────────────────────────────────────────┘
🛠️ Putting It All Together — Your Complete RAG Pipeline in One File
Now you have learned all 7 steps. Let us write a complete pipeline that connects everything — from a user's question to the final AI answer.
This is the full RAG chatbot engine. You pass it a question, it automatically fetches the relevant chunks from Oracle 23ai, sends them with the question to the OCI LLM, and returns the final answer. This is the heart of any GenAI application! ❤️🤖
# ---------------------------------------------------------------
# PURPOSE: The complete GenAI RAG pipeline in one class.
# Input: A user's question (plain English)
# Output: A smart AI-generated answer from YOUR private data
# ---------------------------------------------------------------
import oci
import oracledb
import array
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import (
EmbedTextDetails, ChatDetails, OnDemandServingMode,
CohereChatRequest
)
class OCIGenAIPipeline:
"""
Complete End-to-End GenAI RAG Pipeline on OCI.
Connects Oracle 23ai Vector Search + OCI Generative AI.
"""
def __init__(self, compartment_id: str, db_user: str, db_password: str, db_dsn: str):
# Store configuration
self.compartment_id = compartment_id
self.db_user = db_user
self.db_password = db_password
self.db_dsn = db_dsn
# Connect to OCI services
config = oci.config.from_file("~/.oci/config")
self.ai = GenerativeAiInferenceClient(config)
print("✅ OCI GenAI Pipeline initialized and ready!")
def embed(self, text: str, input_type: str = "SEARCH_QUERY") -> list:
"""Convert text to embedding using OCI Cohere Embed."""
req = EmbedTextDetails(
inputs = [text],
serving_mode = OnDemandServingMode(model_id="cohere.embed-multilingual-v3"),
compartment_id = self.compartment_id,
input_type = input_type
)
response = self.ai.embed_text(req)
return response.data.embeddings[0]
def retrieve(self, question: str, top_k: int = 3) -> list:
"""Find most relevant chunks from Oracle 23ai."""
q_embedding = self.embed(question, "SEARCH_QUERY")
q_array = array.array("f", q_embedding)
conn = oracledb.connect(user=self.db_user, password=self.db_password, dsn=self.db_dsn)
cursor = conn.cursor()
cursor.execute("""
SELECT chunk_text
FROM hospital_knowledge
ORDER BY VECTOR_DISTANCE(embedding, :qv, COSINE) ASC
FETCH FIRST :k ROWS ONLY
""", qv=q_array, k=top_k)
chunks = [row[0] for row in cursor.fetchall()]
cursor.close()
conn.close()
return chunks
def generate(self, question: str, context_chunks: list) -> str:
"""Generate answer using OCI LLM with the retrieved context."""
context = "\n---\n".join(context_chunks)
prompt = f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
chat_req = CohereChatRequest(
message = prompt,
preamble = "You are a helpful assistant. Answer only using the context provided.",
max_tokens = 500,
temperature = 0.1
)
response = self.ai.chat(ChatDetails(
serving_mode = OnDemandServingMode(model_id="cohere.command-r-plus"),
compartment_id = self.compartment_id,
chat_request = chat_req
))
return response.data.chat_response.text
def ask(self, question: str) -> str:
"""The main method — ask a question, get a smart answer!"""
print(f"\n❓ Question: {question}")
print("🔍 Searching knowledge base...")
chunks = self.retrieve(question)
print(f"📌 Found {len(chunks)} relevant chunks")
print("🤖 Generating answer...")
answer = self.generate(question, chunks)
print(f"\n💬 Answer: {answer}")
return answer
# ── RUN IT ──────────────────────────────────────────────────────
pipeline = OCIGenAIPipeline(
compartment_id = "ocid1.compartment.oc1..xxxxx",
db_user = "ADMIN",
db_password = "YourSecurePassword123!",
db_dsn = "your_adb_high"
)
pipeline.ask("What medication was prescribed for high blood pressure?")
pipeline.ask("How many patients were admitted last month?")
pipeline.ask("What are the most common diagnoses in the oncology department?")
⚠️ Common Mistakes Beginners Make (And How to Avoid Them)
Never feed raw, messy PDFs directly into the embedding model. Always clean and chunk first. Garbage in = garbage out. 🗑️
A chunk of 5000 words confuses the LLM — it cannot focus. Keep chunks between 200 and 500 words for best results. Always add a 10–15% overlap between chunks. ✂️
High temperature means the LLM gets "creative" and may invent facts. For RAG applications, always keep temperature between 0.0 and 0.2. 🌡️
Never write your OCI passwords or API keys directly in the Python code. Always use OCI Vault or environment variables to store secrets safely. 🔒
🎯 Quick Reference — OCI Services Used in This Pipeline
- ☁️ OCI Object Storage — Store all raw and processed data files
- 🔄 OCI Data Integration — Move data from multiple sources automatically
- ⚡ OCI Functions — Serverless code to automate pipeline triggers
- 🧠 OCI Generative AI — Embedding models (Cohere) + LLMs (Command R+, Llama)
- 🗄️ Oracle Database 23ai — Vector storage + similarity search (VECTOR data type)
- 🌐 OCI API Gateway — Expose your pipeline as a REST API to your application
- 🔒 OCI Vault — Store passwords and secrets securely
- 📊 OCI Logging & Monitoring — Track pipeline health and performance
Happy building! 🛠️✨
Comments
Post a Comment