Skip to main content

Document Understanding Using OIC AI Capabilities

Calculating read time…

📁 The Problem: 847 Supplier Contracts. One Procurement Manager. No Time.

Anjali is the Procurement Manager at a fast-growing retail company. Her company has 847 active supplier contracts — PDFs, Word documents, scanned paper contracts — stored in a shared drive going back 12 years.

The CFO walks in on a Monday morning:

"Anjali, I need to know — by Thursday — which of our supplier contracts have auto-renewal clauses expiring in the next 90 days, what the penalty clauses are for early termination, and which contracts have price escalation clauses above 5%. Also I need a summary of payment terms across all contracts for the board presentation."

Anjali stares at the screen. 847 contracts. Average 18 pages each. That is 15,246 pages. By Thursday. Manually.

She calls her team. They start reading. Tuesday morning — they have reviewed 43 contracts. 804 left. The CFO already sent a follow-up email.

💡 OCI Document Understanding is the technology that reads those 847 contracts — automatically.

It extracts: key clauses, payment terms, renewal dates, penalty thresholds, party names, contract values — from any PDF or Word document, regardless of format, template, or age. Combined with OIC Gen 3, all 847 contracts are analysed overnight and the extracted data is loaded into Oracle Fusion SCM Procurement — ready for Anjali to review on her dashboard by Tuesday morning. Not Thursday. Tuesday morning.

📑 Section 1: What Is OCI Document Understanding?

Let me explain it the way you would to a 10-year-old.

Imagine you have a very smart friend who has read millions of documents — contracts, invoices, pay slips, tax forms, policies, letters — in their entire life. When you give them a new document, they can instantly:

  • Tell you what type of document it is — "This is a supplier contract, not a purchase order"
  • Find and read every key piece of information — "Payment terms: Net 30. Contract value: ₹45 lakhs. Auto-renewal: Yes, 90 days notice required"
  • Read every single word — even in old scanned documents from 1998
  • Extract all the tables — pricing tables, fee schedules, SLA matrices
  • Understand relationships — "Clause 14.3 modifies the penalties described in Clause 8.1"

OCI Document Understanding is that smart friend — available as a REST API, processing thousands of documents simultaneously, 24/7.

🏛️ OCI Document Understanding vs OCI Vision — What Is the Difference?

This is one of the most common questions. Here is the precise answer:

OCI Vision — optimised for images with structured data. Best for invoice images, receipts, ID cards, product photos, manufacturing defect detection. Treats the document as a visual to extract known fields from.

OCI Document Understanding — optimised for complex multi-page text documents. Best for contracts, policies, reports, research papers, regulatory filings, agreements. Understands document structure, sections, clauses, and relationships between text blocks across pages.

In practice: Supplier invoice → use OCI Vision. Supplier contract → use OCI Document Understanding. Both use the OIC REST Adapter with OCI Signature V1 auth.
📄
Document Classification
Automatically identifies: CONTRACT, INVOICE, BANK_STATEMENT, TAX_FORM, PAYSLIP, RESUME, LEGAL_AGREEMENT, POLICY. No manual labelling needed.
🔑
Key-Value Extraction
Extracts named fields from documents: ContractValue, EffectiveDate, TerminationDate, PaymentTerms, PartyName, AutoRenewalClause, PenaltyPercentage.
🔤
Full Text Extraction (OCR)
Reads every word from scanned PDFs, typed documents, and mixed documents. Preserves page structure, paragraphs, and line positions. Multi-language support.
📊
Table Extraction
Identifies and extracts tables across the document — pricing schedules, SLA matrices, penalty tables, fee structures — as structured row/column data.
🤖
Custom Model (Fine-tuned)
Train a custom extraction model on your specific document templates — your standard contract format, your company's policy documents. Dramatically improves accuracy for domain-specific layouts.
🌍
Multi-Language Support
Works on documents in English, Hindi, Arabic, French, Spanish, German, Japanese, and more. Handles bilingual documents — common in Indian enterprise contracts.

🔌 Section 2: The OCI Document Understanding API

OCI Document Understanding operates differently from OCI Vision. It is primarily designed as an asynchronous job-based API — you submit a job, documents are read from OCI Object Storage, results are written back to Object Storage, and you collect the results. This is perfect for enterprise batch processing.

There is also a synchronous inline API for real-time single-document processing — which we use for on-demand contract review scenarios.

⚡ Synchronous — analyzeDocument
Endpoint: POST /20221101/actions/analyzeDocument

Input: Document as base64 inline OR Object Storage reference
Output: Full extraction result in HTTP response
Best for: Single document, real-time, <20 pages
Latency: 3–15 seconds
🔄 Asynchronous — Processor Jobs
Endpoint: POST /20221101/processorJobs

Input: Object Storage bucket with many documents
Output: JSON result files written to Object Storage
Best for: Batch processing, 100+ documents, multi-page contracts
Latency: Minutes to hours depending on volume

🔌 Synchronous API — analyzeDocument Request Structure

POST https://documentunderstanding.aiservice.{region}.oci.oraclecloud.com/20221101/actions/analyzeDocument

{
  "processorConfig": {
    "processorType": "GENERAL",
    "features": [
      { "featureType": "DOCUMENT_CLASSIFICATION" },
      { "featureType": "KEY_VALUE_EXTRACTION" },
      { "featureType": "TABLE_EXTRACTION" },
      { "featureType": "TEXT_EXTRACTION" }
    ],
    "language": "ENG"
  },
  "document": {
    "source": "INLINE",
    "data": "<base64-encoded-document>",
    "mimeType": "application/pdf"
  },
  "compartmentId": "ocid1.compartment.oc1..{compartment}"
}

📋 Response Structure — what OIC receives back:

{
  "documentMetadata": { "pageCount": 22, "mimeType": "application/pdf" },
  "pages": [
    {
      "pageNumber": 1,
      "words": [/* all words with positions */],
      "lines": [/* text lines */],
      "tables": [/* structured tables found */],
      "documentFields": [/* key-value pairs */]
    }
  ],
  "detectedDocumentTypes": [
    { "documentType": "CONTRACT", "confidence": 0.9821 }
  ],
  "documentClassificationModelVersion": "1.6.0"
}

🏢 Section 3: Enterprise Use Cases — Where Document Understanding Changes Everything

📋
Use Case 1 — Supplier Contract Analysis (Fusion SCM)

Extract: contract value, payment terms, auto-renewal dates, penalty clauses, price escalation terms, governing law. Load into Oracle Fusion Procurement contract repository automatically. Set up alerts for contracts expiring within 90 days. This is what we build in this article.

💰
Use Case 2 — Bank Statement Reconciliation (Fusion Finance)

Extract all transactions from bank statement PDFs — date, description, debit/credit, running balance. Match against Oracle Fusion GL entries. Flag unmatched items for AP/AR team review. Eliminates manual bank rec work for Finance teams.

👤
Use Case 3 — Employee Document Processing (Fusion HCM)

Process employee joining documents — extract: full name, date of birth, qualifications, previous employer, joining date, emergency contact. Auto-populate Oracle Fusion HCM onboarding form. What used to take 45 minutes per employee now takes 8 seconds.

⚖️
Use Case 4 — Legal Document Review & Policy Extraction

Extract key clauses from legal agreements, NDA documents, regulatory filings. Identify obligations, deadlines, liability caps, indemnification clauses. Store structured extracts in Oracle Content Management or ATP database for compliance tracking.

📊
Use Case 5 — Supplier Quotation Processing (Fusion SCM)

Suppliers submit quotations as PDF documents. Extract: item codes, unit prices, lead times, validity periods, discount structures, payment terms. Automatically create Oracle Fusion Purchasing quotation records. Procurement team compares across suppliers without manual data entry.


🏗️ Section 4: Complete Architecture — Supplier Contract Analysis Pipeline

We now build Anjali's solution end to end. The integration receives a supplier contract PDF, analyses it with OCI Document Understanding, extracts all key clauses, and loads the structured data into Oracle Fusion Procurement.

🏗️ Supplier Contract Analysis — Complete OIC Gen 3 Integration Architecture

── DOCUMENT SOURCES ──
☁️ OCI Object Storage
(Contract Inbox Bucket)
🔗 REST Upload
(Supplier Portal)
📧 Email Attachment
(OCI Email Delivery)
⬇️ triggers OIC
⚙️ OIC Gen 3 — CONTRACT_DOCUMENT_ANALYSER Integration
Step 1 — Trigger (REST or OCI Object Storage Event): Receive contractId, supplierNumber, contractObjectName (in OCI Object Storage), documentBase64 (for real-time), businessUnit.
⬇️
Step 2 — Pre-Validation Assign: Check mimeType is PDF or DOCX. Validate base64 length or Object Storage object exists. Set compartmentId, outputBucket, processingStartTime from OIC properties.
⬇️
Step 3 — OCI Document Understanding Invoke (REST Invoke): Call analyzeDocument with features: DOCUMENT_CLASSIFICATION + KEY_VALUE_EXTRACTION + TABLE_EXTRACTION + TEXT_EXTRACTION. Receive full document analysis response including all pages, fields, and tables.
⬇️
Step 4 — Document Type Validation Assign: Verify detectedDocumentType = CONTRACT or LEGAL_AGREEMENT. If not a contract (e.g. someone submitted an invoice), route to exception with clear message. Set documentConfidence variable.
⬇️
Step 5 — Key Field Extraction Assign (XPath from response): Extract: ContractTitle, EffectiveDate, ExpirationDate, AutoRenewalClause, AutoRenewalNoticeDays, ContractValue, Currency, PaymentTerms, PenaltyClause, PenaltyPercentage, PriceEscalationClause, EscalationCap, GoverningLaw, PartyName, SupplierContactName. Calculate daysToExpiry.
⬇️
Step 6 — Business Rules Validation Assign: Check: all required fields present, dates valid, contract value parseable as number, daysToExpiry calculated, escalation cap > 5% flag set, missing field list compiled.
⬇️
Step 7 — Switch: Route by Extraction Quality
✅ HIGH_QUALITY → Auto-load Fusion Procurement contract
⚠️ PARTIAL → Create draft contract + flag for review
❌ LOW_QUALITY → Exception queue + manual processing
⬇️
Step 8 — Oracle Fusion SCM Invoke (Procurement Contract API): Create or update contract record in Oracle Fusion Procurement with all extracted fields. Set alert flags for auto-renewal and expiry notifications.
⬇️
Step 9 — Archive Full Extraction to ATP + Object Storage: Store complete OCI Document Understanding JSON response in OCI Object Storage. Insert structured extracted fields into ATP contract analysis table for reporting.
⬇️
Step 10 — REST Response: Return: { extractionStatus, fusionContractId, extractedFields{}, alertFlags{}, missingFields[], processingTimeMs, confidence }

⚙️ Section 5: Step-by-Step Build in OIC Gen 3

📌 Step 1: Create the Integration

⚙️ Integration Settings

Name:CONTRACT_DOCUMENT_ANALYSER
Style:App Driven Orchestration
Version:01.00.0000
Description:Processes supplier contract PDFs using OCI Document Understanding. Extracts key clauses and terms. Loads structured data into Oracle Fusion Procurement.

📌 Step 2: Configure the REST Trigger

Why: The REST trigger accepts the contract document as base64 for real-time use cases. For batch processing, we use an OCI Event trigger when a file is placed in Object Storage — but the payload structure is the same.

🔌 REST Trigger Configuration

Endpoint Name:analyseContract
Relative URI:/contract/analyse
Method:POST

📋 Request JSON Sample:

{
  "contractId": "CONTRACT-2024-00847",
  "supplierNumber": "SUPP-00234",
  "businessUnit": "IN_CORP",
  "documentBase64": "JVBERi0xLjQKJ...",
  "mimeType": "application/pdf",
  "fileName": "GlobalParts_SupplyAgreement_2024.pdf",
  "submittedBy": "anjali.sharma@company.com",
  "expectedContractValue": 4500000
}

📋 Response JSON Sample:

{
  "status": "SUCCESS",
  "extractionQuality": "HIGH_QUALITY",
  "fusionContractId": "300000012345001",
  "extractedContractTitle": "Master Supply Agreement - GlobalParts Manufacturing",
  "extractedEffectiveDate": "2024-01-15",
  "extractedExpirationDate": "2025-01-14",
  "daysToExpiry": 127,
  "extractedContractValue": "4500000.00",
  "extractedPaymentTerms": "Net 30 days from invoice date",
  "autoRenewalFlag": "YES",
  "autoRenewalNoticeDays": 90,
  "penaltyPercentage": "15.0",
  "priceEscalationCap": "7.5",
  "alertFlags": { "escalationAbove5Pct": true, "highPenalty": true, "renewalWithin90Days": false },
  "overallConfidence": 0.9344,
  "processingTimeMs": 8742,
  "missingFields": []
}

📌 Step 3: Create the OCI Document Understanding Connection

Why: OCI Document Understanding uses the same OCI Signature V1 authentication as OCI Vision and OCI Language. Once you build this connection, you follow the same pattern for every OCI AI service.

🔌 OCI Document Understanding REST Connection

OIC Console → Connections → Create → Adapter: REST
Connection Name:OCI_DOC_UNDERSTANDING_CONN
Base URL:https://documentunderstanding.aiservice.ap-mumbai-1.oci.oraclecloud.com
Security Policy:OCI Signature Version 1
Tenancy OCID:ocid1.tenancy.oc1..{your_tenancy}
User OCID:ocid1.user.oc1..{oic_service_user_ocid}
Private Key (.pem):Upload RSA private key. Same key used for OCI Vision / Language if already set up.
Fingerprint:aa:bb:cc:dd:ee:...
✅ Required IAM Policy: Allow group OICIntegrationGroup to use ai-document-understanding in compartment {YourCompartment}
Click Test → "Connection tested successfully" → Save

📌 Step 4: Pre-Validation Assign Activity

⚙️ Assign Activity — "preValidateContract"

Variable: isDocumentValid
if(string-length(normalize-space($triggerRequest.documentBase64)) > 500 and ($triggerRequest.mimeType = 'application/pdf' or $triggerRequest.mimeType = 'application/vnd.openxmlformats-officedocument.wordprocessingml.document'), 'true', 'false')
Variable: compartmentId
$OICProperties.OCI_COMPARTMENT_ID — configure in OIC Properties panel, never hardcode
Variable: processorType
"GENERAL" — use GENERAL for standard contracts. Change to custom model OCID if you have a fine-tuned model.
Variable: processingStartTimeMs
fn:current-dateTime()

📌 Step 5: OCI Document Understanding REST Invoke

Why: This is the central AI call. We request all four feature types in one API call — classification tells us it is a contract, key-value extraction pulls out the clauses, table extraction finds pricing schedules, and text extraction gives us the full text for any clause not captured by key-value.

📑 REST Invoke — OCI Document Understanding analyzeDocument

Invoke Name:analyseContractDocument
Relative URI:/20221101/actions/analyzeDocument
Method:POST
Timeout (seconds):60 — contracts can be large multi-page PDFs

📋 Request JSON Sample (paste in wizard):

{
  "processorConfig": {
    "processorType": "GENERAL",
    "features": [
      { "featureType": "DOCUMENT_CLASSIFICATION" },
      { "featureType": "KEY_VALUE_EXTRACTION" },
      { "featureType": "TABLE_EXTRACTION" },
      { "featureType": "TEXT_EXTRACTION" }
    ],
    "language": "ENG"
  },
  "document": {
    "source": "INLINE",
    "data": "base64stringhere",
    "mimeType": "application/pdf"
  },
  "compartmentId": "ocid1.compartment.oc1..sample"
}

📋 Response JSON Sample (paste in wizard):

{
  "documentMetadata": { "pageCount": 22, "mimeType": "application/pdf" },
  "detectedDocumentTypes": [
    { "documentType": "CONTRACT", "confidence": 0.98 }
  ],
  "pages": [
    {
      "pageNumber": 1,
      "documentFields": [
        { "fieldLabel": { "name": "ContractTitle" }, "fieldValue": { "text": "Master Supply Agreement", "confidence": 0.96 } },
        { "fieldLabel": { "name": "EffectiveDate" }, "fieldValue": { "text": "January 15, 2024", "confidence": 0.97 } },
        { "fieldLabel": { "name": "ExpirationDate" }, "fieldValue": { "text": "January 14, 2025", "confidence": 0.97 } },
        { "fieldLabel": { "name": "ContractValue" }, "fieldValue": { "text": "INR 45,00,000", "confidence": 0.93 } },
        { "fieldLabel": { "name": "PaymentTerms" }, "fieldValue": { "text": "Net 30 days from invoice date", "confidence": 0.95 } },
        { "fieldLabel": { "name": "AutoRenewal" }, "fieldValue": { "text": "Yes, unless 90 days written notice given", "confidence": 0.91 } },
        { "fieldLabel": { "name": "PenaltyClause" }, "fieldValue": { "text": "15% of remaining contract value for early termination", "confidence": 0.89 } },
        { "fieldLabel": { "name": "PriceEscalation" }, "fieldValue": { "text": "Annual increase not to exceed 7.5%", "confidence": 0.92 } }
      ],
      "tables": [{ "rows": [] }]
    }
  ]
}

📌 Step 6: Key Field Extraction Assign

Why: The OCI Document Understanding response nests fields inside a pages array. We extract each field using XPath predicates — searching the documentFields array across all pages for each label name. We also apply date formatting and numeric cleaning in this step.

⚙️ Assign Activity — "extractContractFields"

These XPath expressions search across all pages of the document. The // operator searches recursively through the entire response tree — useful because different fields may appear on different pages:

Variable: extractedContractTitle
$docResponse.pages/documentFields[fieldLabel/name='ContractTitle']/fieldValue/text
Variable: extractedEffectiveDateRaw
$docResponse.pages/documentFields[fieldLabel/name='EffectiveDate']/fieldValue/text
Raw date string. We convert it to ISO format (2024-01-15) in a subsequent Assign step for Fusion compatibility.
Variable: extractedExpirationDateRaw
$docResponse.pages/documentFields[fieldLabel/name='ExpirationDate']/fieldValue/text
Variable: extractedContractValueRaw
$docResponse.pages/documentFields[fieldLabel/name='ContractValue']/fieldValue/text
Returns "INR 45,00,000" — we clean this with translate() to strip INR, commas, spaces and get a clean decimal number.
Variable: extractedContractValueNumeric
number(translate(translate($extractedContractValueRaw, 'INR ₹,', ''), ' ', ''))
Strips currency symbols, commas, and spaces. Converts to a clean decimal Oracle Fusion can use.
Variable: extractedPaymentTerms
$docResponse.pages/documentFields[fieldLabel/name='PaymentTerms']/fieldValue/text
Variable: autoRenewalText
$docResponse.pages/documentFields[fieldLabel/name='AutoRenewal']/fieldValue/text
Variable: autoRenewalFlag
if(contains(lower-case($autoRenewalText), 'yes') or contains(lower-case($autoRenewalText), 'automatic') or contains(lower-case($autoRenewalText), 'renew'), 'YES', 'NO')
Smart detection — checks for multiple ways a contract can express auto-renewal
Variable: autoRenewalNoticeDays (extract the number from the text)
if(contains($autoRenewalText, '90'), '90', if(contains($autoRenewalText, '60'), '60', if(contains($autoRenewalText, '30'), '30', '0')))
Extracts the notice period in days from the raw clause text
Variable: penaltyText
$docResponse.pages/documentFields[fieldLabel/name='PenaltyClause']/fieldValue/text
Variable: priceEscalationText
$docResponse.pages/documentFields[fieldLabel/name='PriceEscalation']/fieldValue/text
Variable: documentType
$docResponse.detectedDocumentTypes[1]/documentType
Variable: documentConfidence
$docResponse.detectedDocumentTypes[1]/confidence

📌 Step 7: Business Rules & Alert Flag Assign

⚙️ Assign Activity — "applyBusinessRules"

This is where OIC applies your company's specific rules to the extracted data — the intelligence layer that turns raw extracted text into actionable business signals.

Variable: daysToExpiry
days-from-duration(xs:date($extractedExpirationDateFormatted) - fn:current-date())
Calculates exact number of days until contract expires from today
Variable: flagRenewalWithin90Days
if($daysToExpiry <= 90 and $daysToExpiry >= 0, 'true', 'false')
Flags contracts expiring soon — this answers the CFO's first question
Variable: priceEscalationCapNumeric
if(contains($priceEscalationText, '7.5'), 7.5, if(contains($priceEscalationText, '5'), 5.0, if(contains($priceEscalationText, '10'), 10.0, 0)))
Variable: flagEscalationAbove5Pct
if($priceEscalationCapNumeric > 5, 'true', 'false')
Answers the CFO's third question — escalation above 5%
Variable: hasRequiredFields
if(string-length($extractedContractTitle) > 0 and string-length($extractedExpirationDateRaw) > 0 and $extractedContractValueNumeric > 0, 'true', 'false')
Variable: extractionQuality (routing decision)
if($documentConfidence >= 0.90 and $hasRequiredFields = 'true', 'HIGH_QUALITY', if($documentConfidence >= 0.70 and $hasRequiredFields = 'true', 'PARTIAL', 'LOW_QUALITY'))

📌 Step 8: Switch Activity — Route by Extraction Quality

🔀 Switch Activity — Three Routing Branches

✅ HIGH_QUALITY
Condition: $extractionQuality = 'HIGH_QUALITY'

① Call Oracle Fusion SCM Procurement REST API → create Contract record (status=ACTIVE)
② Set all extracted fields as contract attributes
③ Set alert flags for renewal and escalation
④ $actionTaken = "CONTRACT_AUTO_CREATED"
⚠️ PARTIAL
Condition: $extractionQuality = 'PARTIAL'

① Create draft contract in Fusion (status=DRAFT)
② Populate all fields that were successfully extracted
③ Add comment: "AI extraction partial — review fields: [missingFields]"
④ Send OIC notification to procurement team
❌ LOW_QUALITY
Default branch

① No Fusion record created
② Log to exception ATP table with failure details
③ Send Teams notification: "Contract [id] requires manual processing"
④ Archive original document to review bucket

📌 Step 9: Oracle Fusion Procurement Contract Creation (HIGH_QUALITY Branch)

Why this matters: For Anjali's scenario — this step loads all 847 contracts into Oracle Fusion Procurement automatically. Every contract, with its renewal date, value, payment terms, and alert flags, is now searchable in the Fusion Procurement module. The CFO's Thursday report is ready before Tuesday lunch.

🏛️ Oracle Fusion SCM Procurement REST Invoke — Create Contract

Connection:ORACLE_FUSION_SCM_CONN
REST URI:/fscmRestApi/resources/11.13.18.05/purchaseContracts
Method:POST

📋 Fusion Contract Payload (mapped from OIC variables):

{
  "ContractTitle": $extractedContractTitle,
  "ContractNumber": $triggerRequest.contractId,
  "SupplierNumber": $triggerRequest.supplierNumber,
  "BusinessUnit": $triggerRequest.businessUnit,
  "EffectiveDate": $extractedEffectiveDateFormatted,
  "ExpirationDate": $extractedExpirationDateFormatted,
  "ContractAmount": $extractedContractValueNumeric,
  "Currency": "INR",
  "PaymentTerms": $extractedPaymentTerms,
  "Status": "ACTIVE",
  "ContractSource": "OIC_DOCUMENT_AI",
  "DFF_AutoRenewal_c": $autoRenewalFlag,
  "DFF_AutoRenewalNoticeDays_c": number($autoRenewalNoticeDays),
  "DFF_PriceEscalationCap_c": $priceEscalationCapNumeric,
  "DFF_PenaltyClause_c": $penaltyText,
  "DFF_AIConfidenceScore_c": $documentConfidence,
  "DFF_AlertRenewalFlag_c": $flagRenewalWithin90Days,
  "DFF_AlertEscalationFlag_c": $flagEscalationAbove5Pct,
  "DFF_DaysToExpiry_c": $daysToExpiry,
  "Notes": concat("AI-processed by OIC Document Understanding. Document: ", $triggerRequest.fileName, ". Confidence: ", $documentConfidence)
}

📌 Step 10: Archive to ATP Table + Object Storage

Why two storage targets? Object Storage holds the complete raw OCI Document Understanding JSON response — all pages, all word positions, all extracted text — for audit and reprocessing. The ATP table holds just the key structured fields for fast SQL reporting — the CFO's dashboard queries ATP, not Object Storage.

💾 ATP Table — CONTRACT_ANALYSIS_LOG (INSERT via REST DB Adapter)

INSERT INTO CONTRACT_ANALYSIS_LOG (
  CONTRACT_ID, SUPPLIER_NUMBER, BUSINESS_UNIT,
  EXTRACTED_TITLE, EFFECTIVE_DATE, EXPIRATION_DATE,
  CONTRACT_VALUE, PAYMENT_TERMS, AUTO_RENEWAL_FLAG,
  AUTO_RENEWAL_NOTICE_DAYS, PENALTY_TEXT, PRICE_ESCALATION_CAP,
  DAYS_TO_EXPIRY, FLAG_RENEWAL_90_DAYS, FLAG_ESCALATION_ABOVE_5PCT,
  AI_CONFIDENCE_SCORE, EXTRACTION_QUALITY, FUSION_CONTRACT_ID,
  DOCUMENT_OBJECT_NAME, PROCESSED_TIMESTAMP
) VALUES (
  :contractId, :supplierNumber, :businessUnit,
  :extractedTitle, :effectiveDate, :expirationDate,
  :contractValue, :paymentTerms, :autoRenewalFlag,
  :autoRenewalNoticeDays, :penaltyText, :priceEscalationCap,
  :daysToExpiry, :flagRenewal, :flagEscalation,
  :confidence, :extractionQuality, :fusionContractId,
  :documentObjectName, SYSTIMESTAMP
)

🔄 Section 6: Batch Processing — Analysing All 847 Contracts Overnight

For Anjali's 847-contract scenario, we use the OCI Document Understanding asynchronous Processor Job API. This pattern processes documents in parallel at massive scale.

🔄 Async Batch Architecture — 847 Contracts Overnight

Preparation: Upload all 847 contract PDFs to OCI Object Storage bucket contracts-inbox/batch-2024-Q4/ Each file named: CONTRACT-XXXX.pdf. This happens when contracts arrive via email or supplier portal.
⬇️
⏰ OIC Schedule (Sunday 22:00) — CONTRACT_BATCH_INITIATOR integration triggers
Lists all objects in inbox bucket. Calls OCI Document Understanding createProcessorJob: POST /20221101/processorJobs → { inputLocation: {source: OBJECT_STORAGE, namespaceName, bucketName, prefix: "contracts-inbox/batch-2024-Q4/"}, outputLocation: {namespaceName, bucketName: "contracts-results", prefix: "results-2024-Q4/"}, processorConfig: { processorType: "GENERAL", features: [...] } } Receives: processorJobId = "job-contract-2024-Q4-001"
⬇️ OCI processes 847 documents in parallel
OCI Document Understanding writes results to Object Storage
Each contract: contracts-results/results-2024-Q4/CONTRACT-XXXX_result.json
Each JSON file contains: all pages, all documentFields, classification, tables. 847 result files written.
⬇️ OCI Events triggers when job status = SUCCEEDED
OCI Event triggers OIC CONTRACT_BATCH_PROCESSOR integration
OIC lists all result files. For each result JSON:
1. Reads file from Object Storage
2. Runs same extraction + validation + Fusion creation logic as real-time flow
3. Uses OIC parallel-for-each to process up to 50 contracts simultaneously
⬇️
✅ Monday 06:30 AM — Batch Summary Report emailed to Anjali
"Overnight contract analysis complete: 779 contracts loaded to Oracle Fusion Procurement. 52 partial extractions flagged for review. 16 exceptions. Contracts with auto-renewal in 90 days: 34. Contracts with escalation >5%: 127. Contracts expiring this month: 18."
The CFO's board presentation data is ready. It is 6:30 AM Monday. The CFO asked on Monday morning last week. ✅

👤 Section 7: Bonus Use Case — Employee Document Processing in Fusion HCM

The same OIC + OCI Document Understanding integration pattern works for HR onboarding. When a new employee submits their joining documents — offer letter, educational certificates, ID proof, previous employment letter — this integration processes them automatically.

👤 HR Onboarding Document Processing — OIC Flow

📄
Offer Letter
Extract: Job title, Start date, Salary, Grade, Reporting Manager, Location
🎓
Degree Certificate
Extract: Degree, University, Year of passing, Specialisation, Grade/CGPA
🪪
ID Document
Extract: Full name, Date of birth, ID number, Expiry date, Address
💼
Experience Letter
Extract: Previous employer, Role, Tenure, Last drawn salary, Reason for leaving
OIC Action After Extraction: Auto-populate Oracle Fusion HCM new hire record — Personal Details, Qualifications, Previous Employment, and Identity Documents — in a single orchestration that runs in 12 seconds. What previously took the HR coordinator 45 minutes per employee now happens instantly at document submission. For a company hiring 200 people per month, that is 150 hours of HR time saved every month.

🚫 Common Mistakes

🔴
Assuming Fields Always Appear on Page 1

Contract payment terms might be on page 3. Auto-renewal clauses on page 14. If your XPath only searches pages[1]/documentFields, you miss most of the critical clauses. Always use the pages/documentFields path (all pages) or explicitly loop through all page numbers. Use the // deep-search operator sparingly as it impacts performance on large documents.

🔴
Using TEXT_EXTRACTION Alone Instead of KEY_VALUE_EXTRACTION

Some developers only request TEXT_EXTRACTION (raw OCR) and then try to parse contract clauses from raw text using string functions in OIC. This is 10x harder and far less accurate than using KEY_VALUE_EXTRACTION, which does the semantic understanding for you and returns pre-labelled fields with confidence scores. Always use both — KEY_VALUE for structured extraction, TEXT as fallback for searching unextracted clauses.

🔴
Not Handling Date Format Variations

OCI Document Understanding returns dates in the format they appear in the document: "January 15, 2024", "15/01/2024", "15-Jan-24", "2024-01-15". If you pass any of these directly to Oracle Fusion without normalisation, Fusion rejects the payload. Always normalise all extracted dates to ISO 8601 format (YYYY-MM-DD) in an Assign Activity before calling Fusion APIs.

🔴
Processing Scanned Image-Only PDFs Without Verifying OCR Quality

Old contracts scanned at low resolution (below 150 DPI) produce very low confidence scores. Running these through Document Understanding and auto-loading them into Fusion creates garbage data. Add a pre-check: if documentConfidence < 0.70 and pageCount < 3, route to a manual review queue with a message "Low-quality scan — please rescan at 300 DPI and resubmit."


✅ Best Practices

✅ Build a Field Name Synonym Table

Different suppliers use different terminology. "Contract Value" vs "Total Agreement Value" vs "Purchase Commitment". Build an OIC Lookup Table mapping all synonyms to your standard field names. Before running XPath extractions, check the synonym table first. This dramatically improves extraction coverage across diverse contract templates without requiring separate integrations per supplier.

✅ Train a Custom Model for Your Standard Contract Template

If your company uses a standard contract template — and most do — train an OCI Document Understanding custom model on 50–100 examples of your template. A custom model extracts fields from your specific template with 95%+ confidence vs 85–90% for the GENERAL model. The training effort is 2–3 days and the accuracy improvement is dramatic. Use OCI Data Labelling Service to annotate training samples.

✅ Implement a Confidence Waterfall Strategy

For each field: (1) Try KEY_VALUE_EXTRACTION — if confidence > threshold, use it. (2) If confidence is low, try TEXT_EXTRACTION with a targeted regex search through the full document text. (3) If still not found, mark as MISSING and flag for human review. This three-tier approach maximises extraction coverage without sacrificing accuracy for critical fields.

✅ Use OIC Business Identifiers for Contract Analytics

Set OIC Business Identifiers on: contractId, supplierNumber, extractionQuality, daysToExpiry, flagRenewalWithin90Days, flagEscalationAbove5Pct. This transforms OIC monitoring from a technical operations tool into a contract management analytics dashboard. The CFO can filter: "Show me all contracts where escalationAbove5Pct = true processed this quarter" — directly from the OIC monitoring console.


🛡️ Security Considerations

🛡️ Concern Risk in Document Processing ✅ Mitigation in OIC
Contract Data in LLM Training Contract values and terms sent to a public AI API could be used for model training OCI Document Understanding runs entirely within your OCI tenancy. Data never leaves. No public API used.
PII in Employee Documents Employee ID numbers, date of birth, salary — all PII under GDPR/PDPA Mask PII fields in logs. Store in OCI Vault-encrypted Object Storage. Restrict ATP table access to HR role only via OCI IAM data policies.
Unauthorised Contract Submission Any user submitting a contract document triggers Fusion record creation OIC REST trigger must be authenticated. Add pre-check: verify submittedBy email is in approved procurement team group via IDCS/OCI IAM API call before processing.
Document Injection Attack Malicious PDF with embedded scripts or crafted content to exploit the AI extractor Always use OCI Object Storage source (not inline for production). OCI scans objects for malware. Never execute any content from extracted text fields in OIC logic.

🎓 Interview Questions — Senior OIC Developer Level

❓ "What is the difference between OCI Vision and OCI Document Understanding? When do you use each in an OIC integration?"

OCI Vision is optimised for image analysis — structured documents like invoices, receipts, ID cards — where you need to extract known fields from a visual. OCI Document Understanding is optimised for complex multi-page text documents like contracts, policies, and agreements — where semantic understanding of clauses, sections, and relationships matters. In OIC: use Vision for invoice images from suppliers, use Document Understanding for supplier contracts, employment agreements, and policy documents. Both use OCI Signature V1 auth via REST adapter in OIC.

❓ "How do you extract a field like AutoRenewalNoticeDays when OCI Document Understanding returns it as a full sentence in English?"

OCI Document Understanding returns the AutoRenewal field as a complete sentence: "Yes, unless 90 days written notice given prior to expiration." In OIC, I use two approaches: (1) XPath contains() checks — if the text contains '90', extract 90 as the value. This works for common notice periods. (2) For more robust extraction, I pass the raw extracted text through OCI Language API (Key Phrase Extraction) to get the numeric value. The pragmatic OIC approach is to build a nested XPath if-else checking for common notice period values (30, 45, 60, 90, 120) and default to 0 if none match — flagging that contract for manual review of the notice period.

❓ "How do you handle 847 contracts for batch processing in OIC? Walk me through the architecture."

Three-stage architecture: (1) Preparation: all contracts uploaded to OCI Object Storage contracts-inbox/ bucket as they arrive. (2) Submission: OIC Schedule integration (Sunday night) calls OCI Document Understanding createProcessorJob API, pointing to the input bucket and a results output bucket. OCI processes all documents in parallel — typically takes 30–90 minutes for 847 contracts. (3) Collection: OCI Events notification triggers an OIC integration when the processor job reaches SUCCEEDED status. That OIC integration reads all result JSON files from the output bucket, runs business validation and Fusion contract creation for each, using OIC parallel-for-each to process 50 contracts simultaneously. Summary email sent to procurement manager on Monday morning.

❓ "A contract is scanned at low DPI and the AI returns low confidence scores for all fields. How does your OIC integration handle this?"

Three layers of handling: (1) Pre-call check: if base64 length suggests very small file size for a multi-page document, flag as potentially low-quality before even calling Document Understanding. (2) Post-call check: if documentConfidence < 0.70 overall, route to LOW_QUALITY branch — no Fusion record created. (3) Field-level check: even if document confidence is acceptable, if a critical field (ContractValue, ExpirationDate) has individual confidence < 0.80, add that field to the missingFields list and set extractionQuality = PARTIAL. The integration never silently loads low-confidence data into Fusion — it routes to manual review with a clear message: "Low confidence on ExpirationDate — please verify manually."


🎉 Final Summary

📑 OCI Document Understanding: Classifies, extracts key-value pairs, reads tables, and OCRs full text from any PDF or Word document. Best for complex multi-page text documents — contracts, policies, agreements, bank statements.
🔗 OIC Connection: REST Adapter with OCI Signature Version 1. Same auth pattern as OCI Vision and OCI Language. Configure once — reuse across all OCI AI services.
⚙️ Key OIC Components: Pre-validation Assign → OCI DU REST Invoke → Field Extraction Assign (XPath predicates) → Business Rules Assign (date calculation, numeric normalisation, flag setting) → Switch routing → Fusion Invoke → ATP Insert + Object Storage Archive.
🔑 Critical XPath Pattern: $docResponse.pages/documentFields[fieldLabel/name='FieldName']/fieldValue/text — searches across all pages for the named field. Always extract both text value and confidence score for every field.
🔄 Batch Processing: createProcessorJob API processes 847+ documents overnight. OCI Events triggers OIC when complete. Results in Object Storage are processed with parallel-for-each in OIC for speed.
🛡️ Data Stays in OCI Tenancy: All document processing happens inside your Oracle Cloud tenancy boundary. Supplier contracts and employee documents never reach any external AI API.
📊 Business Impact: 847 contracts analysed overnight → CFO's contract data ready Monday morning. 200 employee onboarding documents per month → 150 hours of HR manual work eliminated. Contract renewal risk visibility → zero missed auto-renewals.
🌱 The Broader Lesson:

OCI Document Understanding, OCI Vision, and OCI Language are three AI services that each solve a different aspect of the same fundamental challenge: unstructured data is everywhere, and structured systems like Oracle Fusion can only work with structured data.

OIC is the integration layer that connects these AI services to your Oracle ecosystem. Every JSON payload transformation you build in OIC, every XPath extraction, every Switch routing decision — these are the bridges that turn AI output into Oracle business transactions.

You are not just building integrations. You are building intelligence pipelines that make your Oracle Fusion investment exponentially more valuable.


Read Every Document. Extract Every Insight. Connect Every System. 📑 ⚙️

Comments