Skip to main content

Retrieval-Augmented Generation using OIC

Calculating read time…

"What Is Our Policy on This?" — The Question That Kills Productivity

It is 4:47 PM on a Thursday. Arjun, a Procurement Analyst, is processing a supplier invoice for ₹78 lakhs. The supplier has sent two separate invoices instead of one consolidated invoice — a common tactic to stay below the CFO's single-invoice approval threshold of ₹50 lakhs.

Arjun suspects this might be invoice splitting — a policy violation. But he is not sure. He searches the shared drive. He finds 11 different policy documents. He checks the Procurement Policy v3.2. He checks the Finance Controls Manual. He checks the Supplier Code of Conduct. Forty-five minutes later — he still cannot find a definitive answer about invoice splitting policy. He escalates to his manager. His manager checks with Legal. The invoice sits unpaid for three more days.

What if Arjun could just type: "Can a supplier split one purchase order into multiple invoices to stay below the CFO approval threshold?" — and get an answer in 3 seconds, with the exact policy clause cited, ready to use?


That is RAG — Retrieval-Augmented Generation. And it is the most powerful pattern in enterprise AI.

It combines the precision of a search engine (finds the right document, the right clause, the right page) with the intelligence of a language model (understands the question, reads the clause, formulates a clear answer in plain English, cites the source). Built inside OIC Gen 3 — it connects directly to your Oracle Fusion applications, your policy knowledge base, and your enterprise data — all within OCI.

💡 What You Will Build in This Article:

The Oracle Procurement Policy Q&A System — a complete RAG pipeline using OIC Gen 3 that:
① Accepts a natural language question from an Oracle user
② Converts it to a vector embedding using OCI Generative AI
③ Searches OCI OpenSearch (vector database) for relevant policy chunks
④ Sends retrieved context + question to OCI GenAI LLM
⑤ Returns a grounded, cited answer with source document references
⑥ Logs the Q&A to Oracle Fusion for audit and compliance tracking

Full implementation — knowledge base setup — indexing pipeline — query pipeline — every OIC step explained.

📚 Section 1: What Is RAG — Explained Like You Are 10

Imagine you have a very smart friend who has a terrible memory — they cannot remember anything after their school days. But they have access to the world's best library and can read any book in 1 second.

When you ask them a question, they do two things:

  1. Search the library for the most relevant books and pages on your question
  2. Read those pages and give you an intelligent, well-explained answer based on what they just read — not from memory

That is RAG. The "library" is your knowledge base. The "search" is vector search. The "reading and answering" is the LLM. The LLM stops making things up because it has real, current, authoritative documents to ground its answer on.

❌ Pure LLM (No RAG)
Arjun asks: "What is our policy on invoice splitting?"

LLM answers from training data — which does not include YOUR company's internal policies. It gives a generic, plausible-sounding answer that may be completely wrong for your organisation. Result: Hallucination. Dangerous for compliance decisions.
⚠️ Pure Keyword Search
Arjun searches for "invoice splitting" in the policy portal.

Gets 23 documents. Must read each one. Cannot ask follow-up questions. Cannot get a synthesised answer across multiple policy sections. Result: Information found but not understood. Back to square one.
✅ RAG (Search + Generate)
Arjun asks: "Can suppliers split one PO into multiple invoices to bypass CFO threshold?"

RAG finds the exact policy clause. LLM reads it. Answers: "No. Procurement Policy v3.2, Section 4.1 prohibits invoice splitting. Intentional splitting to bypass approval thresholds is a Level 2 policy violation." Result: Accurate, cited, grounded answer in 3 seconds.
💡 Why RAG Matters So Much for Oracle Enterprise Applications:

A pure LLM does not know: your approval thresholds, your supplier blacklists, your payment terms, your internal controls, your implementation-specific configurations, or any policy change made after its training cutoff.

RAG gives the LLM access to your knowledge — not generic internet knowledge. It answers based on your Procurement Policy v3.2. Your Finance Controls Manual. Your Oracle Fusion implementation guide. Your company's specific rules. This is the difference between a useful enterprise tool and an expensive autocomplete that confidently gives wrong answers.

🔢 Section 2: Vector Embeddings — The Magic Behind Semantic Search

To understand RAG, you must understand vector embeddings. This is the part most tutorials skip or explain badly. We will explain it simply and completely.

🎯 The Problem with Keyword Search

Traditional keyword search looks for exact word matches. Arjun searches for "invoice splitting" — it finds documents containing those exact words. But what if the policy document says "fragmented billing" or "distributed invoicing" or "split billing across multiple tax periods"? Keyword search misses all of these — even though they mean the same thing.

Vector search understands meaning, not just words.

🔢 What Is a Vector Embedding?

An embedding model reads a piece of text and converts it into a list of numbers — a "vector" — where similar meanings produce mathematically similar vectors.

🔢 How Embeddings Capture Meaning — A Simple Illustration

Text 1: Arjun's Question
"Can suppliers split one PO into multiple invoices to bypass approval threshold?"
Vector: [0.82, -0.14, 0.67, 0.33, -0.55, 0.91, ...]
1024 numbers representing the meaning of this sentence
≈
Text 2: Policy Clause (HIGH MATCH)
"Fragmented billing across multiple invoices from a single PO to circumvent financial controls is prohibited."
Vector: [0.79, -0.18, 0.71, 0.28, -0.51, 0.88, ...]
Similarity score: 0.94 ← VERY CLOSE
Same Question
"Can suppliers split one PO into multiple invoices to bypass approval threshold?"
Vector: [0.82, -0.14, 0.67, 0.33, -0.55, 0.91, ...]
≠
Text 3: Unrelated Clause (LOW MATCH)
"Employees must submit travel expense claims within 30 days of the journey date."
Vector: [0.12, 0.74, -0.33, 0.91, 0.22, -0.14, ...]
Similarity score: 0.11 ← VERY FAR
💡 The embedding model captured that "split invoices + bypass approval" is semantically close to "fragmented billing + circumvent controls" — even though they share zero keywords. This is why vector search finds relevant information that keyword search completely misses.

📐 Cosine Similarity — How We Measure "How Close Are Two Meanings?"

Two vectors are compared using cosine similarity — a measure of how similar two vectors are. The result is a number between -1 and 1.

0.90 – 1.00
Highly Relevant
Same topic, same intent. Return as top result.
0.70 – 0.89
Relevant
Related topic. Include in context for LLM.
0.50 – 0.69
Loosely Related
May be useful as supplemental context. Include with caution.
0.00 – 0.49
Not Relevant
Discard. Different topic entirely.
🏛️ Architect Insight — Choosing the Right Similarity Threshold:

In OCI OpenSearch, you set a minimum similarity score for results. Start with 0.70 as your threshold. Too high (0.90+) → you miss relevant documents with slightly different wording. Too low (0.50) → you include noise that confuses the LLM with unrelated context. 0.70–0.75 is the sweet spot for enterprise policy Q&A. Tune based on your specific knowledge base and question types after initial testing.

🏗️ Section 3: Complete RAG Architecture — Two Pipelines

Every RAG system has two distinct pipelines. Understanding both is essential before touching OIC.

🏗️ Pipeline 1 — INGESTION (Run Once / On Update)
"Loading your knowledge base into the vector database"

1. Read source documents (policy PDFs, manuals, guides)
2. Split into small, overlapping chunks (~500 words each)
3. Embed each chunk using OCI Generative AI Embeddings
4. Store chunk text + vector + metadata in OCI OpenSearch

Run when: new documents added, policies updated. Not on every query.
🔍 Pipeline 2 — QUERY (Run on Every Question)
"Answering a question using the knowledge base"

1. Receive user's question
2. Embed the question using the same embedding model
3. Search OCI OpenSearch for most similar chunks
4. Send retrieved chunks + question to LLM
5. Return grounded, cited answer

Run on: every user question. Typically completes in 2–5 seconds.

🏛️ Oracle RAG Architecture — Ingestion + Query Pipelines

── PIPELINE 1: INGESTION (OIC DOCUMENT_INDEXER Integration) ──
📄 Policy PDFs
OCI Object Storage
→
✂️ OIC
Chunk Splitter
→
🔢 OCI GenAI
Embed API
cohere-embed-v3
→
🗄️ OCI OpenSearch
Vector Index
knn_vector field
── PIPELINE 2: QUERY (OIC POLICY_RAG_ENGINE Integration) ──
👤 User Question
Oracle Fusion / REST
→
🔢 OCI GenAI
Embed Question
same model
→
🔍 OCI OpenSearch
k-NN Vector Search
top-5 chunks
→
🧠 OCI GenAI LLM
Generate Answer
with context
→
✅ Grounded
Cited Answer
Both pipelines use OIC REST Adapter with OCI Signature V1 auth. The same OCI_GENAI_CONN and OCI_OPENSEARCH_CONN connections serve both pipelines. OCI OpenSearch stores the chunks as both text (for returning to LLM) and vectors (for similarity search).

📚 Section 4: Setting Up the Knowledge Base — What to Index and How

Before writing a single OIC step, you must decide what goes into the knowledge base. This is a content architecture decision that determines how accurate and useful your RAG system will be.

📂 Step 4A: Identify Your Source Documents

📋
Tier 1 — Critical Policy Documents (Must Index)
Procurement Policy Manual, Finance Controls Framework, Supplier Code of Conduct, Approval Authority Matrix, Expense Policy, Travel Policy. These are the documents users ask about most often. Every version matters — tag with version number and effective date.
📘
Tier 2 — Oracle Fusion Implementation Guides (High Value)
Your organisation-specific Oracle Fusion setup documentation: approval workflow configurations, business unit structures, chart of accounts, supplier numbering conventions, payment terms catalogue. Users frequently ask "how is X configured in our Fusion system?"
📄
Tier 3 — Process SOPs and Training Materials (Useful)
How-to guides for Oracle Fusion tasks, training manuals, frequently asked questions from the help desk, audit findings with resolutions. These answer procedural questions: "How do I create a blanket purchase agreement in Oracle Fusion SCM?"

✂️ Step 4B: The Chunking Strategy — Splitting Documents for Vector Search

Why chunking matters: You cannot embed an entire 80-page policy manual as one vector. The vector would represent the "average meaning" of 80 pages — useless for finding specific clauses. You must split the document into small chunks where each chunk represents one coherent idea or policy statement. The quality of your chunking strategy determines the quality of your search results.

✂️ Chunking Strategy Comparison

Strategy Chunk Size Best For Overlap
Fixed Size 500 tokens General policy documents, consistent formatting 50 tokens
Sentence-Based 3–5 sentences Dense legal text where each sentence is a standalone rule 1 sentence
Section-Based ✅ Recommended One policy section Policy manuals with numbered sections and headings Section header
Hierarchical Parent + child chunks Complex documents with section hierarchy (manuals) Parent context
💡 Recommended for Oracle Policy Documents: Section-Based chunking with 50-token overlap. Each chunk = one numbered policy section (e.g. "4.1 Invoice Splitting Policy") + the section header. The overlap ensures that the context from the section title is included in each chunk, improving search relevance. Store the section number and document name as metadata on each chunk.

🗄️ Step 4C: OCI OpenSearch Index Design

⚙️ OCI OpenSearch Index Mapping — policy_knowledge_base

This is the schema of your vector database index. Each document chunk is stored as one OpenSearch document with these fields. This mapping is created once via the OpenSearch REST API before ingestion begins.

PUT /policy_knowledge_base
{
  "settings": {
    "index.knn": true,
    "index.knn.algo_param.ef_search": 512
  },
  "mappings": {
    "properties": {
      "chunkId": { "type": "keyword" },
      "chunkText": { "type": "text" },
      "chunkVector": {
        "type": "knn_vector",
        "dimension": 1024,
        "method": { "name": "hnsw", "space_type": "cosinesimil", "engine": "nmslib" }
      },
      "documentName": { "type": "keyword" },
      "documentVersion": { "type": "keyword" },
      "sectionNumber": { "type": "keyword" },
      "sectionTitle": { "type": "text" },
      "documentCategory": { "type": "keyword" },
      "effectiveDate": { "type": "date" },
      "pageNumber": { "type": "integer" }
    }
  }
}
💡 Key fields explained:
chunkVector — the 1024-dimensional embedding vector. dimension=1024 matches the Cohere embed-v3 model output.
hnsw — Hierarchical Navigable Small World. The algorithm used for approximate nearest-neighbour search. Fast and highly accurate.
cosinesimil — cosine similarity as the distance metric. Measures angle between vectors (semantic similarity).
chunkText — the original text that gets returned to the LLM as context.

⚙️ Section 5: Building Pipeline 1 — DOCUMENT_INDEXER Integration

This OIC integration is triggered whenever a new policy document is uploaded or updated. It reads the document, splits it into chunks, embeds each chunk, and loads them into OCI OpenSearch. Run it once initially for all existing documents, then automatically for every future update.

⚙️ OIC DOCUMENT_INDEXER — Ingestion Pipeline Steps

Step 1 — OCI Events Trigger (Object Storage PUT Event): Triggered when a new file is uploaded to policy-documents/ bucket. Receives: namespace, bucketName, objectName, eventTime. Also supports manual REST trigger for initial bulk load.
⬇️
Step 2 — Read Document from OCI Object Storage (REST Invoke): GET /n/{namespace}/b/policy-documents/o/{objectName}. Receives document content as base64. Extract document metadata: filename, version, effective date (parsed from filename convention: PolicyName_vX.X_YYYYMMDD.pdf).
⬇️
Step 3 — Extract Text via OCI Document Understanding (REST Invoke): Call analyzeDocument with TEXT_EXTRACTION feature. Receive full text per page. Concatenate all page text into one document string with page markers: [PAGE 1]\n{text}\n[PAGE 2]\n{text}...
⬇️
Step 4 — Chunk the Document (JavaScript Action in OIC): Split full document text into chunks. Strategy: split on section headings (regex: /\n\d+\.\d+\s+[A-Z]/). Each chunk = section header + section content. Max chunk size: 600 words. If section > 600 words: split at paragraph boundaries. Store as array of chunk objects: { chunkId, sectionNumber, sectionTitle, chunkText, pageNumber }.
⬇️ For each chunk (parallel-for-each)
Step 5 — Embed Each Chunk (OCI GenAI Embeddings API): POST to /20231130/actions/embedText. Model: cohere.embed-english-v3.0. Input type: SEARCH_DOCUMENT. Receive 1024-dimensional float array. This is the vector for this chunk.
⬇️
Step 6 — Index Chunk in OCI OpenSearch (REST Invoke): POST to OpenSearch index. Document: { chunkId, chunkText, chunkVector (the 1024 floats), documentName, documentVersion, sectionNumber, sectionTitle, documentCategory, effectiveDate, pageNumber }. Response: index confirmation with _id.
⬇️ After all chunks indexed
Step 7 — Log Indexing Completion to ATP: INSERT into DOCUMENT_INDEX_LOG: documentName, version, totalChunks, indexedChunks, indexingTimeMs, indexedAt, indexedBy. This gives you a complete audit of what is in your knowledge base and when it was last updated.

📌 OCI Generative AI Embeddings API — Exact Configuration

🔢 REST Invoke — embedText API

Connection:OCI_GENAI_CONN (same connection as chat API)
Relative URI:/20231130/actions/embedText
Method:POST

📋 Request Sample (for INGESTION — indexing a chunk):

{
  "compartmentId": "ocid1.compartment.oc1..{compartment}",
  "servingMode": {
    "servingType": "ON_DEMAND",
    "modelId": "cohere.embed-english-v3.0"
  },
  "inputs": ["4.1 Invoice Splitting Policy\nThe practice of dividing a single purchase..."],
  "inputType": "SEARCH_DOCUMENT"
}

📋 Response Sample:

{
  "embeddings": [
    [0.8241, -0.1342, 0.6677, 0.3291, -0.5512, 0.9103, ... 1018 more floats ...]
  ],
  "modelId": "cohere.embed-english-v3.0"
}
⚠️ Critical: inputType must match usage. Use "SEARCH_DOCUMENT" when embedding chunks for indexing (ingestion pipeline). Use "SEARCH_QUERY" when embedding the user's question (query pipeline). Using the wrong inputType significantly degrades search accuracy. The Cohere model is optimised differently for document content vs short queries.

🔍 Section 6: Building Pipeline 2 — POLICY_RAG_ENGINE Integration

This is the integration that answers Arjun's question in real-time. Every time a user asks a policy question, this integration runs: embed the question → search OpenSearch → build context → call LLM → return cited answer.

⚙️ OIC POLICY_RAG_ENGINE — Query Pipeline Steps

Step 1 — REST Trigger: Receive: question (natural language), userId, userRole, fusionModuleContext (e.g. "PROCUREMENT"), sessionId. Return: { answer, sources[], confidence, sessionId }.
⬇️
Step 2 — Pre-Processing Assign: Validate question not empty and length > 5 chars. Clean question text (normalize-space). Build moduleFilter based on fusionModuleContext (PROCUREMENT → filter to Procurement + Finance policy docs). Set topK=5, minSimilarity=0.70.
⬇️
Step 3 — Embed the User Question (OCI GenAI Embed API): POST to embedText. inputType: "SEARCH_QUERY" (NOT SEARCH_DOCUMENT — this is critical). Input: user's question text. Receive 1024-dimensional question vector. Store as $questionVector.
⬇️
Step 4 — Vector Search in OCI OpenSearch (REST Invoke): POST k-NN search query to OpenSearch. Pass $questionVector. Request top 5 most similar chunks. Filter by documentCategory matching fusionModuleContext. Receive: chunk texts + metadata + similarity scores.
⬇️
Step 5 — Build RAG Context Assign: Assemble the retrieved chunks into a numbered context block. Filter out any chunk with score < 0.70. Format as: "SOURCE 1 [ProcurementPolicy_v3.2, Section 4.1]: {chunkText}\nSOURCE 2 [FinanceControls_v2.1, Section 8.3]: {chunkText}..."
⬇️
Step 6 — Build RAG Prompt Assign: Construct the full LLM prompt: system preamble (Oracle analyst persona) + instruction (answer ONLY from provided context, cite sources, JSON output format) + retrieved context block + user question.
⬇️
Step 7 — Call OCI GenAI LLM with Context (REST Invoke): POST to /20231130/actions/chat. Model: cohere.command-r-plus. Temperature: 0.1. MaxTokens: 800. The LLM reads the policy chunks and generates a grounded, cited answer.
⬇️
Step 8 — Parse Response + Extract Sources Assign: Extract answer text and sources from JSON response. Build sources array from the retrieved chunks metadata (documentName, sectionNumber, version, pageNumber). Calculate confidence score from top chunk similarity score.
⬇️
Step 9 — Log Q&A to ATP (Optional but Recommended): INSERT into POLICY_QA_LOG: sessionId, userId, userRole, question, answer, sources, confidence, retrievedChunkIds[], processingTimeMs, timestamp. Critical for audit, analytics, and model improvement.
⬇️
Step 10 — REST Response: Return: { answer, sources[], confidence, processingTimeMs, sessionId, disclaimer: "This answer is based on indexed policy documents. Always verify with your compliance team for binding decisions." }

📌 OCI OpenSearch — k-NN Vector Search Query (Step 4 Detail)

🔍 OpenSearch REST Invoke — k-NN Similarity Search

Connection:OCI_OPENSEARCH_CONN
Base URL:https://{your-opensearch-cluster}.opensearch.{region}.oci.oraclecloud.com
Relative URI:/policy_knowledge_base/_search
Method:POST
Auth:Basic Auth (OpenSearch username/password stored in OCI Vault)

📋 k-NN Search Request Body (mapped in OIC Data Mapper):

{
  "size": 5,
  "query": {
    "bool": {
      "must": [
        {
          "knn": {
            "chunkVector": {
              "vector": $questionVector,
              "k": 5
            }
          }
        }
      ],
      "filter": [
        {
          "terms": {
            "documentCategory": [$moduleFilter]
          }
        }
      ]
    }
  },
  "_source": ["chunkText", "documentName", "documentVersion", "sectionNumber", "sectionTitle", "pageNumber"],
  "min_score": 0.70
}

📋 OpenSearch Response (what OIC receives):

{
  "hits": {
    "total": { "value": 3 },
    "hits": [
      {
        "_score": 0.9421,
        "_source": {
          "chunkText": "4.1 Invoice Splitting Policy\nThe practice of dividing a single purchase order into multiple invoices...",
          "documentName": "Procurement_Policy_v3.2",
          "sectionNumber": "4.1",
          "sectionTitle": "Invoice Splitting Policy",
          "pageNumber": 23
        }
      },
      // ... 2 more hits ...
    ]
  }
}

📌 The RAG Prompt — Context + Question + Citation Instructions

⚙️ Assign Activity — "buildRAGPrompt"

System Preamble:
"You are an expert Oracle enterprise policy advisor. You answer questions ONLY using the policy document excerpts provided to you. You never use general knowledge or make assumptions. If the provided excerpts do not contain enough information to answer the question, say exactly: 'The provided policy documents do not contain sufficient information to answer this question. Please consult your compliance team.' Always cite your sources using [Document Name, Section X.X]."

User Message (built with concat in OIC):
concat(
  "Answer the following question using ONLY the policy excerpts below. ",
  "Respond in valid JSON with keys: answer (string, max 200 words), ",
  "sources (array of objects: {document, section, pageNumber}), ",
  "confidence (HIGH/MEDIUM/LOW based on how directly the excerpts address the question), ",
  "caveat (string - any important limitation or 'null' if none). ",
  "\n\nPOLICY EXCERPTS:\n",
  $assembledContext,
  "\n\nQUESTION: ", $triggerRequest.question
)

🎯 Real RAG Answer — What Arjun Gets

✅ RAG-Generated Answer — Grounded and Cited

Arjun's Question: "Can suppliers split one purchase order into multiple invoices to bypass the CFO approval threshold?"

Answer: No. Invoice splitting to circumvent approval thresholds is explicitly prohibited. Procurement Policy v3.2, Section 4.1 states that dividing a single purchase order into multiple invoices to avoid financial controls is a Level 2 policy violation. This applies regardless of whether the intent to circumvent controls can be proven — the pattern itself triggers the violation classification. Consequences include mandatory referral to the Internal Audit team and potential suspension of the supplier. Finance Controls Framework v2.1, Section 8.3 further requires that invoices above ₹50L receive CFO approval and that all invoices from the same PO within a 30-day window are aggregated for threshold calculation purposes.

Sources:
• Procurement_Policy_v3.2, Section 4.1 — Invoice Splitting Policy (Page 23) [Confidence: 0.94]
• Finance_Controls_v2.1, Section 8.3 — Approval Thresholds & Aggregation Rules (Page 47) [Confidence: 0.87]

Confidence: HIGH
⚠️ This answer is based on indexed policy documents. Always verify with your compliance team for binding decisions.

📡 Section 7: Complete Trigger and Response Payloads

🔌 REST Trigger — POLICY_RAG_ENGINE

Relative URI:/policy/ask
Method:POST

📋 Request JSON Sample:

{
  "question": "Can suppliers split one PO into multiple invoices to bypass CFO threshold?",
  "userId": "arjun.sharma@company.com",
  "userRole": "PROCUREMENT_ANALYST",
  "fusionModuleContext": "PROCUREMENT",
  "sessionId": "sess-2024-09-16-001",
  "maxSources": 5,
  "minConfidenceThreshold": 0.70
}

📋 Response JSON Sample:

{
  "status": "SUCCESS",
  "answer": "No. Invoice splitting is prohibited under Procurement Policy v3.2 Section 4.1...",
  "confidence": "HIGH",
  "sources": [
    { "document": "Procurement_Policy_v3.2", "section": "4.1", "pageNumber": 23, "similarity": 0.94 },
    { "document": "Finance_Controls_v2.1", "section": "8.3", "pageNumber": 47, "similarity": 0.87 }
  ],
  "caveat": "null",
  "sessionId": "sess-2024-09-16-001",
  "processingTimeMs": 2847,
  "disclaimer": "AI-generated answer based on indexed policy documents. Verify with compliance team for binding decisions."
}

🏛️ Section 8: Embedding RAG Inside Oracle Fusion Applications

The real power of building RAG in OIC Gen 3 is that you can embed it directly inside Oracle Fusion workflows. Here are three high-value integration patterns:

🔗 Pattern 1 — Oracle Fusion Procurement: Inline Policy Check During PO Creation

When a Procurement user raises a Purchase Order in Oracle Fusion, a custom button "Check Policy" calls the OIC RAG endpoint with context from the PO (supplier name, amount, category). The RAG system automatically queries: "What approval is needed for this PO type and amount?" and returns the exact policy, approval authority required, and any special conditions — inline in the Fusion UI. No tab-switching, no manual policy lookup.

OIC Components: Oracle Fusion REST REST Adapter (trigger) → POLICY_RAG_ENGINE (sub-integration call) → Response displayed in Fusion flex field / descriptive text
🔗 Pattern 2 — Oracle CX Service: Agent Assist for Customer Service Representatives

A customer service agent receives a complaint about a warranty claim. As they type notes in Oracle CX, an OIC integration monitors the SR and automatically calls the RAG system with the complaint context. The system retrieves the relevant warranty policy section and suggests a response — pre-populated in the CX reply field. The agent reviews, edits if needed, and sends. Average handle time drops by 40%.

OIC Components: Oracle CX Business Events (SR update event) → POLICY_RAG_ENGINE → Update SR Notes via CX REST API
🔗 Pattern 3 — Oracle Fusion HCM: Employee Self-Service Policy Q&A

Employees ask HR policy questions through an Oracle HCM embedded chatbot widget. Questions about leave entitlements, expense policies, performance review timelines, and remote work policies are answered instantly using the RAG system. The HR team's query volume drops by 60%. The RAG log provides analytics on which policies generate the most questions — revealing which policies need clarification or better communication.

OIC Components: Oracle HCM REST trigger (chatbot webhook) → POLICY_RAG_ENGINE (HR category filter) → Response to HCM chatbot widget

🚫 Common Mistakes

🔴
Using SEARCH_DOCUMENT inputType for Query Embeddings

The Cohere embed model is asymmetric — it uses different internal representations for document content vs queries. Using SEARCH_DOCUMENT for both ingestion AND query embeddings causes poor search accuracy (similarity scores 20–30% lower than optimal). Always use SEARCH_DOCUMENT when embedding policy chunks (ingestion) and SEARCH_QUERY when embedding user questions (query pipeline). This single mistake is the most common cause of "RAG finds unrelated documents."

🔴
Chunks Too Large — One Chunk Per Document

Embedding an entire 80-page policy manual as one chunk means the vector represents the "average" of 80 pages — which matches nothing precisely. When Arjun asks about invoice splitting, the vector search finds the manual but cannot pinpoint the specific section. Split documents into sections of 300–600 tokens maximum. Each chunk must represent one coherent policy statement — not a whole document.

🔴
Not Updating the Index When Policies Change

Your Procurement Policy v3.2 is updated to v3.3 — the invoice splitting threshold changes from ₹50L to ₹75L. The old chunks are still in OpenSearch. The RAG system now gives the old (wrong) answer. Always trigger the ingestion pipeline automatically when a new document version is uploaded to Object Storage. Add a document version filter to queries: only return chunks from the latest version of each document. Never let outdated policy chunks co-exist with current ones in the same index.

🔴
Telling the LLM to Answer From Its Own Knowledge When Context Is Insufficient

Some RAG prompts say: "If the provided context is insufficient, use your general knowledge to answer." This completely defeats the purpose of RAG. The LLM's general knowledge does not know your specific policies. It will confidently hallucinate a plausible-sounding answer that may be dangerous for compliance decisions. The correct instruction: "If the provided context does not contain sufficient information, respond: 'The policy documents do not address this question. Please consult your compliance team.'"


✅ Best Practices

✅ Add Metadata Filters to Narrow Search Before Vector Comparison

Vector search across all 5,000 policy chunks is slower and noisier than searching within the right category. When Arjun's question comes from the Procurement module, filter to documentCategory IN ["PROCUREMENT","FINANCE"] before running the k-NN search. This reduces the search space from 5,000 chunks to ~800 — faster results and higher precision because irrelevant HR or Legal policy chunks are excluded from candidates.

✅ Implement Hybrid Search — Vector + Keyword Together

Vector search is excellent for semantic similarity. Keyword search is excellent for exact terms (regulation numbers, specific clause references like "Section 4.1"). Combine both in OpenSearch using the bool query with a should clause for keyword match alongside the knn clause. Weight the results: 70% vector similarity + 30% keyword match. This hybrid approach outperforms pure vector search by 15–25% on enterprise policy documents which contain specific terminology users reference exactly.

✅ Build a Human Feedback Loop Into the RAG Log

After each RAG answer, let users rate it: 👍 Helpful / 👎 Not Helpful / ✏️ Incorrect. Store ratings in the POLICY_QA_LOG ATP table. After 200 ratings: run a query — which questions consistently get thumbs-down? These reveal: (a) policies that need better chunking, (b) policies missing from the knowledge base, (c) questions where vector search finds irrelevant chunks. Feedback-driven knowledge base improvement is what separates a demo RAG from a production-grade enterprise system.

✅ Cache Embeddings for Common Questions

The most-asked questions in a Policy Q&A system tend to repeat — "What is the travel expense limit?", "What is the approval authority for invoices above ₹10L?". These questions will be asked hundreds of times. Store the question embedding (vector) and its top search results in OCI Cache (Redis-compatible) with a 24h TTL. Cache hit: skip embedding API call + skip vector search. Serve from cache instantly. This reduces latency from ~3 seconds to ~0.3 seconds for cached questions and cuts embedding API costs by 40%+ on high-volume deployments.


🎓 Interview Questions — RAG Architect Level

❓ "What is RAG and why is it needed even when you have a powerful LLM?"

RAG (Retrieval-Augmented Generation) is a pattern that gives the LLM access to your specific knowledge at query time — rather than relying on general training data. It is needed because: (1) LLMs have a training cutoff — they do not know your 2024 policy updates. (2) LLMs do not know your organisation-specific rules — your approval thresholds, your supplier codes, your Oracle Fusion configuration. (3) Pure LLMs hallucinate on specific factual questions — they give plausible-sounding but potentially wrong answers. RAG prevents hallucination by giving the LLM real, current, authoritative documents as grounding context. The LLM's job shifts from "know the answer" to "read these documents and formulate an answer" — which it does very well.

❓ "Explain vector embeddings and cosine similarity to someone who has never heard of them."

A vector embedding converts text into a list of numbers (a vector) that represents the meaning of the text mathematically. Texts with similar meanings produce similar vectors — even if they use different words. For example, "invoice splitting" and "fragmented billing to bypass controls" produce vectors that are mathematically close, even though they share no words. Cosine similarity measures how close two vectors are: a score near 1.0 means the meanings are very similar; near 0.0 means completely unrelated. In a RAG system, we convert the user's question to a vector, then search the knowledge base for stored text chunks whose vectors are most similar. This finds relevant policy sections based on meaning — not keyword matching — which is dramatically more accurate for natural language questions.

❓ "How do you build a RAG ingestion pipeline in OIC Gen 3?"

Five steps: (1) OCI Events trigger fires when a policy document is uploaded to OCI Object Storage. (2) OIC reads the document and calls OCI Document Understanding to extract full text. (3) A JavaScript action in OIC splits the extracted text into section-based chunks (~500 tokens each, 50-token overlap). (4) For each chunk, OIC calls the OCI GenAI embedText API with inputType=SEARCH_DOCUMENT to get a 1024-dimensional vector. (5) OIC calls OCI OpenSearch REST API to index the chunk with its vector and metadata (document name, version, section number, page number). The OpenSearch index has a knn_vector field with cosinesimil distance metric. The whole pipeline is triggered automatically on document upload — the knowledge base is always current.

❓ "What chunking strategy do you recommend for Oracle policy documents and why?"

Section-based chunking — splitting at numbered section headings (4.1, 4.2, 4.3) — with a maximum of 500 tokens per chunk and 50-token overlap. The reasoning: policy documents are organised into numbered sections where each section is a coherent, standalone policy statement. Splitting at section boundaries preserves the semantic coherence of each chunk — the vector for section 4.1 accurately represents "Invoice Splitting Policy." Fixed-size splitting might cut a section in half, producing chunks whose vectors represent partial, incoherent concepts that match poorly in search. The section number and title are prepended to every chunk as context — so even if the body is short, the vector encodes the full policy context including the heading.

❓ "How would you ensure the RAG system stays current when Oracle Fusion policies are updated?"

Three-layer currency management: (1) Automatic re-indexing via OCI Events: whenever a document is updated in the OCI Object Storage policy bucket, the DOCUMENT_INDEXER integration automatically triggers, deletes old chunks for that document (using OpenSearch bulk delete by documentName + old version), and re-indexes the new version with the updated version number. (2) Version filtering in queries: all vector search queries include a filter for documentVersion = "latest" — preventing old chunks from surfacing even if deletion was delayed. (3) ATP index log monitoring: an OIC scheduled integration runs daily, checks the DOCUMENT_INDEX_LOG for documents not updated in 90+ days, and sends an alert to the Knowledge Base Administrator: "These policy documents have not been reviewed in 90 days — please verify they are current." This prevents stale policies from silently persisting in the knowledge base.


🎉 Final Summary — RAG in OIC Gen 3 at a Glance

🔢 Vector Embeddings: Text converted to 1024-dimensional number arrays where similar meanings = similar vectors. OCI GenAI cohere-embed-v3 generates them. Use SEARCH_DOCUMENT for indexing, SEARCH_QUERY for questions — never mix them.
🗄️ OCI OpenSearch: Your vector database. knn_vector field with hnsw algorithm and cosinesimil distance. Index every policy chunk with metadata. k-NN search returns top-5 most similar chunks. Filter by documentCategory before search for precision and speed.
🏗️ Two Pipelines: DOCUMENT_INDEXER (runs when policies update — chunk → embed → index). POLICY_RAG_ENGINE (runs on every question — embed question → vector search → build context → LLM → cited answer). Both built as OIC Gen 3 integrations.
✍️ The RAG Prompt: System preamble (answer ONLY from provided context, never general knowledge) + assembled context (retrieved chunks numbered with source labels) + user question + JSON output format. The grounding instruction is non-negotiable.
✅ Grounded, Cited Answers: Every RAG answer includes: the answer text, source document names + section numbers + page numbers, confidence level, and a disclaimer. Arjun gets his policy answer in 3 seconds with the exact clause cited — not a generic hallucinated response.
🏛️ Fusion Integration Patterns: Inline policy check during PO creation, agent assist in Oracle CX, employee self-service in HCM. All powered by the same POLICY_RAG_ENGINE OIC integration with different module context filters.
🔄 Continuous Improvement: User feedback ratings → identify weak spots. Automatic re-indexing on policy update → always current. Question caching → sub-second responses for repeated queries. ATP Q&A log → analytics on what users ask most.
🌱 The Transformation RAG Delivers:

Before RAG: Arjun spends 45 minutes searching policy documents. He escalates. The invoice sits unpaid for 3 days. Three people are involved in answering one compliance question.

After RAG: Arjun types his question in Oracle Fusion. 3 seconds later he has the exact policy, the clause number, the consequence of violation, and a confidence level. He makes the decision himself. Correctly. The invoice is either processed or flagged for investigation — immediately.

Multiply this across 200 Procurement Analysts, 500 Finance users, and 3,000 employees asking HR policy questions. RAG does not just save time — it democratises expert policy knowledge to every person in your Oracle ecosystem, at the moment they need it.


Search with Precision. Retrieve with Context. Answer with Confidence. 🔍 🧠 📚

Comments