"What Is Our Policy on This?" — The Question That Kills Productivity
It is 4:47 PM on a Thursday. Arjun, a Procurement Analyst, is processing a supplier invoice for ₹78 lakhs. The supplier has sent two separate invoices instead of one consolidated invoice — a common tactic to stay below the CFO's single-invoice approval threshold of ₹50 lakhs.
Arjun suspects this might be invoice splitting — a policy violation. But he is not sure. He searches the shared drive. He finds 11 different policy documents. He checks the Procurement Policy v3.2. He checks the Finance Controls Manual. He checks the Supplier Code of Conduct. Forty-five minutes later — he still cannot find a definitive answer about invoice splitting policy. He escalates to his manager. His manager checks with Legal. The invoice sits unpaid for three more days.
That is RAG — Retrieval-Augmented Generation. And it is the most powerful pattern in enterprise AI.
It combines the precision of a search engine (finds the right document, the right clause, the right page) with the intelligence of a language model (understands the question, reads the clause, formulates a clear answer in plain English, cites the source). Built inside OIC Gen 3 — it connects directly to your Oracle Fusion applications, your policy knowledge base, and your enterprise data — all within OCI.
The Oracle Procurement Policy Q&A System — a complete RAG pipeline using OIC Gen 3 that:
① Accepts a natural language question from an Oracle user
② Converts it to a vector embedding using OCI Generative AI
③ Searches OCI OpenSearch (vector database) for relevant policy chunks
④ Sends retrieved context + question to OCI GenAI LLM
⑤ Returns a grounded, cited answer with source document references
⑥ Logs the Q&A to Oracle Fusion for audit and compliance tracking
Full implementation — knowledge base setup — indexing pipeline — query pipeline — every OIC step explained.
📚 Section 1: What Is RAG — Explained Like You Are 10
Imagine you have a very smart friend who has a terrible memory — they cannot remember anything after their school days. But they have access to the world's best library and can read any book in 1 second.
When you ask them a question, they do two things:
- Search the library for the most relevant books and pages on your question
- Read those pages and give you an intelligent, well-explained answer based on what they just read — not from memory
That is RAG. The "library" is your knowledge base. The "search" is vector search. The "reading and answering" is the LLM. The LLM stops making things up because it has real, current, authoritative documents to ground its answer on.
LLM answers from training data — which does not include YOUR company's internal policies. It gives a generic, plausible-sounding answer that may be completely wrong for your organisation. Result: Hallucination. Dangerous for compliance decisions.
Gets 23 documents. Must read each one. Cannot ask follow-up questions. Cannot get a synthesised answer across multiple policy sections. Result: Information found but not understood. Back to square one.
RAG finds the exact policy clause. LLM reads it. Answers: "No. Procurement Policy v3.2, Section 4.1 prohibits invoice splitting. Intentional splitting to bypass approval thresholds is a Level 2 policy violation." Result: Accurate, cited, grounded answer in 3 seconds.
A pure LLM does not know: your approval thresholds, your supplier blacklists, your payment terms, your internal controls, your implementation-specific configurations, or any policy change made after its training cutoff.
RAG gives the LLM access to your knowledge — not generic internet knowledge. It answers based on your Procurement Policy v3.2. Your Finance Controls Manual. Your Oracle Fusion implementation guide. Your company's specific rules. This is the difference between a useful enterprise tool and an expensive autocomplete that confidently gives wrong answers.
🔢 Section 2: Vector Embeddings — The Magic Behind Semantic Search
To understand RAG, you must understand vector embeddings. This is the part most tutorials skip or explain badly. We will explain it simply and completely.
🎯 The Problem with Keyword Search
Traditional keyword search looks for exact word matches. Arjun searches for "invoice splitting" — it finds documents containing those exact words. But what if the policy document says "fragmented billing" or "distributed invoicing" or "split billing across multiple tax periods"? Keyword search misses all of these — even though they mean the same thing.
Vector search understands meaning, not just words.
🔢 What Is a Vector Embedding?
An embedding model reads a piece of text and converts it into a list of numbers — a "vector" — where similar meanings produce mathematically similar vectors.
🔢 How Embeddings Capture Meaning — A Simple Illustration
📐 Cosine Similarity — How We Measure "How Close Are Two Meanings?"
Two vectors are compared using cosine similarity — a measure of how similar two vectors are. The result is a number between -1 and 1.
In OCI OpenSearch, you set a minimum similarity score for results. Start with 0.70 as your threshold. Too high (0.90+) → you miss relevant documents with slightly different wording. Too low (0.50) → you include noise that confuses the LLM with unrelated context. 0.70–0.75 is the sweet spot for enterprise policy Q&A. Tune based on your specific knowledge base and question types after initial testing.
🏗️ Section 3: Complete RAG Architecture — Two Pipelines
Every RAG system has two distinct pipelines. Understanding both is essential before touching OIC.
1. Read source documents (policy PDFs, manuals, guides)
2. Split into small, overlapping chunks (~500 words each)
3. Embed each chunk using OCI Generative AI Embeddings
4. Store chunk text + vector + metadata in OCI OpenSearch
Run when: new documents added, policies updated. Not on every query.
1. Receive user's question
2. Embed the question using the same embedding model
3. Search OCI OpenSearch for most similar chunks
4. Send retrieved chunks + question to LLM
5. Return grounded, cited answer
Run on: every user question. Typically completes in 2–5 seconds.
🏛️ Oracle RAG Architecture — Ingestion + Query Pipelines
OCI Object Storage
Chunk Splitter
Vector Index
knn_vector field
Oracle Fusion / REST
k-NN Vector Search
top-5 chunks
Generate Answer
with context
Cited Answer
📚 Section 4: Setting Up the Knowledge Base — What to Index and How
Before writing a single OIC step, you must decide what goes into the knowledge base. This is a content architecture decision that determines how accurate and useful your RAG system will be.
📂 Step 4A: Identify Your Source Documents
✂️ Step 4B: The Chunking Strategy — Splitting Documents for Vector Search
✂️ Chunking Strategy Comparison
| Strategy | Chunk Size | Best For | Overlap |
|---|---|---|---|
| Fixed Size | 500 tokens | General policy documents, consistent formatting | 50 tokens |
| Sentence-Based | 3–5 sentences | Dense legal text where each sentence is a standalone rule | 1 sentence |
| Section-Based ✅ Recommended | One policy section | Policy manuals with numbered sections and headings | Section header |
| Hierarchical | Parent + child chunks | Complex documents with section hierarchy (manuals) | Parent context |
🗄️ Step 4C: OCI OpenSearch Index Design
⚙️ OCI OpenSearch Index Mapping — policy_knowledge_base
This is the schema of your vector database index. Each document chunk is stored as one OpenSearch document with these fields. This mapping is created once via the OpenSearch REST API before ingestion begins.
{
"settings": {
"index.knn": true,
"index.knn.algo_param.ef_search": 512
},
"mappings": {
"properties": {
"chunkId": { "type": "keyword" },
"chunkText": { "type": "text" },
"chunkVector": {
"type": "knn_vector",
"dimension": 1024,
"method": { "name": "hnsw", "space_type": "cosinesimil", "engine": "nmslib" }
},
"documentName": { "type": "keyword" },
"documentVersion": { "type": "keyword" },
"sectionNumber": { "type": "keyword" },
"sectionTitle": { "type": "text" },
"documentCategory": { "type": "keyword" },
"effectiveDate": { "type": "date" },
"pageNumber": { "type": "integer" }
}
}
}
chunkVector — the 1024-dimensional embedding vector. dimension=1024 matches the Cohere embed-v3 model output.hnsw — Hierarchical Navigable Small World. The algorithm used for approximate nearest-neighbour search. Fast and highly accurate.cosinesimil — cosine similarity as the distance metric. Measures angle between vectors (semantic similarity).chunkText — the original text that gets returned to the LLM as context.
⚙️ Section 5: Building Pipeline 1 — DOCUMENT_INDEXER Integration
This OIC integration is triggered whenever a new policy document is uploaded or updated. It reads the document, splits it into chunks, embeds each chunk, and loads them into OCI OpenSearch. Run it once initially for all existing documents, then automatically for every future update.
⚙️ OIC DOCUMENT_INDEXER — Ingestion Pipeline Steps
policy-documents/ bucket. Receives: namespace, bucketName, objectName, eventTime. Also supports manual REST trigger for initial bulk load.
/n/{namespace}/b/policy-documents/o/{objectName}. Receives document content as base64. Extract document metadata: filename, version, effective date (parsed from filename convention: PolicyName_vX.X_YYYYMMDD.pdf).
[PAGE 1]\n{text}\n[PAGE 2]\n{text}...
/\n\d+\.\d+\s+[A-Z]/). Each chunk = section header + section content. Max chunk size: 600 words. If section > 600 words: split at paragraph boundaries. Store as array of chunk objects: { chunkId, sectionNumber, sectionTitle, chunkText, pageNumber }.
📌 OCI Generative AI Embeddings API — Exact Configuration
🔢 REST Invoke — embedText API
OCI_GENAI_CONN (same connection as chat API)
/20231130/actions/embedText
POST
📋 Request Sample (for INGESTION — indexing a chunk):
"compartmentId": "ocid1.compartment.oc1..{compartment}",
"servingMode": {
"servingType": "ON_DEMAND",
"modelId": "cohere.embed-english-v3.0"
},
"inputs": ["4.1 Invoice Splitting Policy\nThe practice of dividing a single purchase..."],
"inputType": "SEARCH_DOCUMENT"
}
📋 Response Sample:
"embeddings": [
[0.8241, -0.1342, 0.6677, 0.3291, -0.5512, 0.9103, ... 1018 more floats ...]
],
"modelId": "cohere.embed-english-v3.0"
}
"SEARCH_DOCUMENT" when embedding chunks for indexing (ingestion pipeline).
Use "SEARCH_QUERY" when embedding the user's question (query pipeline).
Using the wrong inputType significantly degrades search accuracy. The Cohere model is optimised differently for document content vs short queries.
🔍 Section 6: Building Pipeline 2 — POLICY_RAG_ENGINE Integration
This is the integration that answers Arjun's question in real-time. Every time a user asks a policy question, this integration runs: embed the question → search OpenSearch → build context → call LLM → return cited answer.
⚙️ OIC POLICY_RAG_ENGINE — Query Pipeline Steps
📌 OCI OpenSearch — k-NN Vector Search Query (Step 4 Detail)
🔍 OpenSearch REST Invoke — k-NN Similarity Search
OCI_OPENSEARCH_CONN
https://{your-opensearch-cluster}.opensearch.{region}.oci.oraclecloud.com
/policy_knowledge_base/_search
POST
Basic Auth (OpenSearch username/password stored in OCI Vault)
📋 k-NN Search Request Body (mapped in OIC Data Mapper):
"size": 5,
"query": {
"bool": {
"must": [
{
"knn": {
"chunkVector": {
"vector": $questionVector,
"k": 5
}
}
}
],
"filter": [
{
"terms": {
"documentCategory": [$moduleFilter]
}
}
]
}
},
"_source": ["chunkText", "documentName", "documentVersion", "sectionNumber", "sectionTitle", "pageNumber"],
"min_score": 0.70
}
📋 OpenSearch Response (what OIC receives):
"hits": {
"total": { "value": 3 },
"hits": [
{
"_score": 0.9421,
"_source": {
"chunkText": "4.1 Invoice Splitting Policy\nThe practice of dividing a single purchase order into multiple invoices...",
"documentName": "Procurement_Policy_v3.2",
"sectionNumber": "4.1",
"sectionTitle": "Invoice Splitting Policy",
"pageNumber": 23
}
},
// ... 2 more hits ...
]
}
}
📌 The RAG Prompt — Context + Question + Citation Instructions
⚙️ Assign Activity — "buildRAGPrompt"
"You are an expert Oracle enterprise policy advisor. You answer questions ONLY using the policy document excerpts provided to you. You never use general knowledge or make assumptions. If the provided excerpts do not contain enough information to answer the question, say exactly: 'The provided policy documents do not contain sufficient information to answer this question. Please consult your compliance team.' Always cite your sources using [Document Name, Section X.X]."
User Message (built with concat in OIC):
concat(
"Answer the following question using ONLY the policy excerpts below. ",
"Respond in valid JSON with keys: answer (string, max 200 words), ",
"sources (array of objects: {document, section, pageNumber}), ",
"confidence (HIGH/MEDIUM/LOW based on how directly the excerpts address the question), ",
"caveat (string - any important limitation or 'null' if none). ",
"\n\nPOLICY EXCERPTS:\n",
$assembledContext,
"\n\nQUESTION: ", $triggerRequest.question
)
🎯 Real RAG Answer — What Arjun Gets
✅ RAG-Generated Answer — Grounded and Cited
Arjun's Question: "Can suppliers split one purchase order into multiple invoices to bypass the CFO approval threshold?"
Sources:
• Procurement_Policy_v3.2, Section 4.1 — Invoice Splitting Policy (Page 23) [Confidence: 0.94]
• Finance_Controls_v2.1, Section 8.3 — Approval Thresholds & Aggregation Rules (Page 47) [Confidence: 0.87]
Confidence: HIGH
⚠️ This answer is based on indexed policy documents. Always verify with your compliance team for binding decisions.
📡 Section 7: Complete Trigger and Response Payloads
🔌 REST Trigger — POLICY_RAG_ENGINE
/policy/ask
POST
📋 Request JSON Sample:
"question": "Can suppliers split one PO into multiple invoices to bypass CFO threshold?",
"userId": "arjun.sharma@company.com",
"userRole": "PROCUREMENT_ANALYST",
"fusionModuleContext": "PROCUREMENT",
"sessionId": "sess-2024-09-16-001",
"maxSources": 5,
"minConfidenceThreshold": 0.70
}
📋 Response JSON Sample:
"status": "SUCCESS",
"answer": "No. Invoice splitting is prohibited under Procurement Policy v3.2 Section 4.1...",
"confidence": "HIGH",
"sources": [
{ "document": "Procurement_Policy_v3.2", "section": "4.1", "pageNumber": 23, "similarity": 0.94 },
{ "document": "Finance_Controls_v2.1", "section": "8.3", "pageNumber": 47, "similarity": 0.87 }
],
"caveat": "null",
"sessionId": "sess-2024-09-16-001",
"processingTimeMs": 2847,
"disclaimer": "AI-generated answer based on indexed policy documents. Verify with compliance team for binding decisions."
}
🏛️ Section 8: Embedding RAG Inside Oracle Fusion Applications
The real power of building RAG in OIC Gen 3 is that you can embed it directly inside Oracle Fusion workflows. Here are three high-value integration patterns:
When a Procurement user raises a Purchase Order in Oracle Fusion, a custom button "Check Policy" calls the OIC RAG endpoint with context from the PO (supplier name, amount, category). The RAG system automatically queries: "What approval is needed for this PO type and amount?" and returns the exact policy, approval authority required, and any special conditions — inline in the Fusion UI. No tab-switching, no manual policy lookup.
A customer service agent receives a complaint about a warranty claim. As they type notes in Oracle CX, an OIC integration monitors the SR and automatically calls the RAG system with the complaint context. The system retrieves the relevant warranty policy section and suggests a response — pre-populated in the CX reply field. The agent reviews, edits if needed, and sends. Average handle time drops by 40%.
Employees ask HR policy questions through an Oracle HCM embedded chatbot widget. Questions about leave entitlements, expense policies, performance review timelines, and remote work policies are answered instantly using the RAG system. The HR team's query volume drops by 60%. The RAG log provides analytics on which policies generate the most questions — revealing which policies need clarification or better communication.
🚫 Common Mistakes
The Cohere embed model is asymmetric — it uses different internal representations for document content vs queries. Using SEARCH_DOCUMENT for both ingestion AND query embeddings causes poor search accuracy (similarity scores 20–30% lower than optimal). Always use SEARCH_DOCUMENT when embedding policy chunks (ingestion) and SEARCH_QUERY when embedding user questions (query pipeline). This single mistake is the most common cause of "RAG finds unrelated documents."
Embedding an entire 80-page policy manual as one chunk means the vector represents the "average" of 80 pages — which matches nothing precisely. When Arjun asks about invoice splitting, the vector search finds the manual but cannot pinpoint the specific section. Split documents into sections of 300–600 tokens maximum. Each chunk must represent one coherent policy statement — not a whole document.
Your Procurement Policy v3.2 is updated to v3.3 — the invoice splitting threshold changes from ₹50L to ₹75L. The old chunks are still in OpenSearch. The RAG system now gives the old (wrong) answer. Always trigger the ingestion pipeline automatically when a new document version is uploaded to Object Storage. Add a document version filter to queries: only return chunks from the latest version of each document. Never let outdated policy chunks co-exist with current ones in the same index.
Some RAG prompts say: "If the provided context is insufficient, use your general knowledge to answer." This completely defeats the purpose of RAG. The LLM's general knowledge does not know your specific policies. It will confidently hallucinate a plausible-sounding answer that may be dangerous for compliance decisions. The correct instruction: "If the provided context does not contain sufficient information, respond: 'The policy documents do not address this question. Please consult your compliance team.'"
✅ Best Practices
Vector search across all 5,000 policy chunks is slower and noisier than searching within the right category. When Arjun's question comes from the Procurement module, filter to documentCategory IN ["PROCUREMENT","FINANCE"] before running the k-NN search. This reduces the search space from 5,000 chunks to ~800 — faster results and higher precision because irrelevant HR or Legal policy chunks are excluded from candidates.
Vector search is excellent for semantic similarity. Keyword search is excellent for exact terms (regulation numbers, specific clause references like "Section 4.1"). Combine both in OpenSearch using the bool query with a should clause for keyword match alongside the knn clause. Weight the results: 70% vector similarity + 30% keyword match. This hybrid approach outperforms pure vector search by 15–25% on enterprise policy documents which contain specific terminology users reference exactly.
After each RAG answer, let users rate it: 👍 Helpful / 👎 Not Helpful / ✏️ Incorrect. Store ratings in the POLICY_QA_LOG ATP table. After 200 ratings: run a query — which questions consistently get thumbs-down? These reveal: (a) policies that need better chunking, (b) policies missing from the knowledge base, (c) questions where vector search finds irrelevant chunks. Feedback-driven knowledge base improvement is what separates a demo RAG from a production-grade enterprise system.
The most-asked questions in a Policy Q&A system tend to repeat — "What is the travel expense limit?", "What is the approval authority for invoices above ₹10L?". These questions will be asked hundreds of times. Store the question embedding (vector) and its top search results in OCI Cache (Redis-compatible) with a 24h TTL. Cache hit: skip embedding API call + skip vector search. Serve from cache instantly. This reduces latency from ~3 seconds to ~0.3 seconds for cached questions and cuts embedding API costs by 40%+ on high-volume deployments.
🎓 Interview Questions — RAG Architect Level
RAG (Retrieval-Augmented Generation) is a pattern that gives the LLM access to your specific knowledge at query time — rather than relying on general training data. It is needed because: (1) LLMs have a training cutoff — they do not know your 2024 policy updates. (2) LLMs do not know your organisation-specific rules — your approval thresholds, your supplier codes, your Oracle Fusion configuration. (3) Pure LLMs hallucinate on specific factual questions — they give plausible-sounding but potentially wrong answers. RAG prevents hallucination by giving the LLM real, current, authoritative documents as grounding context. The LLM's job shifts from "know the answer" to "read these documents and formulate an answer" — which it does very well.
A vector embedding converts text into a list of numbers (a vector) that represents the meaning of the text mathematically. Texts with similar meanings produce similar vectors — even if they use different words. For example, "invoice splitting" and "fragmented billing to bypass controls" produce vectors that are mathematically close, even though they share no words. Cosine similarity measures how close two vectors are: a score near 1.0 means the meanings are very similar; near 0.0 means completely unrelated. In a RAG system, we convert the user's question to a vector, then search the knowledge base for stored text chunks whose vectors are most similar. This finds relevant policy sections based on meaning — not keyword matching — which is dramatically more accurate for natural language questions.
Five steps: (1) OCI Events trigger fires when a policy document is uploaded to OCI Object Storage. (2) OIC reads the document and calls OCI Document Understanding to extract full text. (3) A JavaScript action in OIC splits the extracted text into section-based chunks (~500 tokens each, 50-token overlap). (4) For each chunk, OIC calls the OCI GenAI embedText API with inputType=SEARCH_DOCUMENT to get a 1024-dimensional vector. (5) OIC calls OCI OpenSearch REST API to index the chunk with its vector and metadata (document name, version, section number, page number). The OpenSearch index has a knn_vector field with cosinesimil distance metric. The whole pipeline is triggered automatically on document upload — the knowledge base is always current.
Section-based chunking — splitting at numbered section headings (4.1, 4.2, 4.3) — with a maximum of 500 tokens per chunk and 50-token overlap. The reasoning: policy documents are organised into numbered sections where each section is a coherent, standalone policy statement. Splitting at section boundaries preserves the semantic coherence of each chunk — the vector for section 4.1 accurately represents "Invoice Splitting Policy." Fixed-size splitting might cut a section in half, producing chunks whose vectors represent partial, incoherent concepts that match poorly in search. The section number and title are prepended to every chunk as context — so even if the body is short, the vector encodes the full policy context including the heading.
Three-layer currency management: (1) Automatic re-indexing via OCI Events: whenever a document is updated in the OCI Object Storage policy bucket, the DOCUMENT_INDEXER integration automatically triggers, deletes old chunks for that document (using OpenSearch bulk delete by documentName + old version), and re-indexes the new version with the updated version number. (2) Version filtering in queries: all vector search queries include a filter for documentVersion = "latest" — preventing old chunks from surfacing even if deletion was delayed. (3) ATP index log monitoring: an OIC scheduled integration runs daily, checks the DOCUMENT_INDEX_LOG for documents not updated in 90+ days, and sends an alert to the Knowledge Base Administrator: "These policy documents have not been reviewed in 90 days — please verify they are current." This prevents stale policies from silently persisting in the knowledge base.
🎉 Final Summary — RAG in OIC Gen 3 at a Glance
Before RAG: Arjun spends 45 minutes searching policy documents. He escalates. The invoice sits unpaid for 3 days. Three people are involved in answering one compliance question.
After RAG: Arjun types his question in Oracle Fusion. 3 seconds later he has the exact policy, the clause number, the consequence of violation, and a confidence level. He makes the decision himself. Correctly. The invoice is either processed or flagged for investigation — immediately.
Multiply this across 200 Procurement Analysts, 500 Finance users, and 3,000 employees asking HR policy questions. RAG does not just save time — it democratises expert policy knowledge to every person in your Oracle ecosystem, at the moment they need it.
Search with Precision. Retrieve with Context. Answer with Confidence. 🔍 🧠 📚
Comments
Post a Comment