Skip to main content

OCI Document Understanding: AI-Powered Document Analysis and Data Extraction

Calculating read time…

Imagine you have a huge pile of invoices, receipts, and forms sitting on your desk. Reading each one by hand would take forever!  OCI Document Understanding is like hiring a super-smart robot that can read any document, pull out the important info, and hand it back to you in seconds. 





📚 What is OCI Document Understanding?

OCI Document Understanding is an AI service on Oracle Cloud Infrastructure (OCI) that can automatically read, understand, and extract information from scanned documents, PDFs, and images — without you writing any complex logic.

Think of it like this 👇

  • You take a photo of your electricity bill 📸
  • You send it to OCI Document Understanding
  • It reads the bill and tells you: "Due amount: ₹1,200. Due date: 15 May."
  • No human needed! ✅
💡 Real-World Analogy:

Think of OCI Document Understanding as a very fast librarian who can read any book, any receipt, any form — and immediately find exactly what you are looking for, in any language, without ever getting tired! 📚

🌟 What Can OCI Document Understanding Do?

OCI Document Understanding has several superpowers. Let's meet them all:

  • 📝 Text Extraction (OCR) — Reads all the text from any scanned image or PDF
  • 🗂️ Document Classification — Automatically figures out what type of document it is (invoice? passport? resume?)
  • 🔑 Key-Value Extraction — Pulls out specific fields like "Invoice Number", "Total Amount", "Date"
  • 📊 Table Extraction — Reads tables inside documents and gives you clean, structured data
  • 🧾 Receipt / Invoice Processing — Specialized understanding for financial documents
  • 🪪 Identity Document Processing — Reads passports, driving licences, ID cards
  • 🎓 Custom Model Training — Train the AI on YOUR own documents for even better results!
✅ DO Remember:

OCI Document Understanding works best with clear, high-resolution images and properly scanned PDFs. The better the document quality, the better the AI result! 🖼️

🗺️ Big Picture — How Does It All Work?

Before we dive into code, let's understand the full journey of a document through OCI:

  ┌───────────────┐     ┌──────────────────────┐     ┌─────────────────────────┐
  │  Your Document│────►│  OCI Object Storage  │────►│  Document Understanding │
  │ (PDF / Image) │     │  (safe cloud folder) │     │     AI Service          │
  └───────────────┘     └──────────────────────┘     └────────────┬────────────┘
                                                                   │
                                                      ┌────────────▼────────────┐
                                                      │    Extracted Results    │
                                                      │  Text · Tables · KV ·  │
                                                      │  Classification etc.   │
                                                      └─────────────────────────┘

Simple flow: Upload document → AI reads it → Get smart results. 🎯

🏗️ Step 1 — Set Up Your OCI Environment

Before we can use OCI Document Understanding, we need to prepare our environment. Think of this like preparing your kitchen before you start cooking! 🍳

Step 1a: Create an OCI Free Tier Account

  • Go to cloud.oracle.com
  • Click "Start for free"
  • Enter your name, email, and set a password
  • Verify your email ✅
  • Done! You now have access to Oracle Cloud 

Step 1b: Install the OCI Python SDK

📌 What is the OCI SDK?

SDK stands for Software Development Kit. Think of it as a toolbox that lets your Python code talk to Oracle Cloud services. Without this toolbox, your code has no way to reach OCI! 🧰
📌 What does the command below do?

It tells your computer: "Hey, please download and install the OCI toolkit so my Python code can use Oracle Cloud." Run it once in your terminal and you are ready to go!
pip install oci

Step 1c: Configure Your OCI Credentials

OCI needs to know who you are before it lets you in. Think of it like showing your ID card at the security gate. 🪪

📌 What does the command below do?

It starts a setup wizard that asks you a few questions (your OCI region, user ID, etc.) and creates a private key file on your computer. OCI will use this key file every time you make a request — like your personal doorkey! 🔑
oci setup config

Follow the prompts. It will create a config file at ~/.oci/config. That's it! ✅

❌ DON'T do this:

Never share your ~/.oci/config file or private key with anyone. It's like sharing your bank PIN — keep it completely secret! 🔒

☁️ Step 2 — Upload Your Document to OCI Object Storage

OCI Document Understanding needs your document to be stored in OCI Object Storage first. Think of Object Storage like a giant Google Drive on Oracle Cloud — a safe place to keep your files before processing. ☁️

Step 2a: Create a Bucket (via Console)

  • Log in to your OCI Console at cloud.oracle.com
  • Go to Storage → Object Storage → Buckets
  • Click "Create Bucket"
  • Give it a name like my-documents-bucket
  • Click "Create" ✅
💡 What is a Bucket?

A bucket is just a folder in the cloud! Just like you put files inside a folder on your computer, you put documents inside a bucket in OCI. 🪣

Step 2b: Upload a File Using Python

📌 What does the code below do?

This code picks up a PDF file from your computer (for example, an invoice) and uploads it to your OCI bucket in the cloud. Think of it like attaching a file to an email, but the file goes to Oracle Cloud instead of someone's inbox! 📤
import oci

# Load your OCI configuration from the default config file
config = oci.config.from_file()

# Create a client to talk to Object Storage
object_storage_client = oci.object_storage.ObjectStorageClient(config)

# Get the namespace (a unique ID for your OCI account's storage area)
namespace = object_storage_client.get_namespace().data

# Name of the bucket we created earlier
bucket_name = "my-documents-bucket"

# Path to the local file we want to upload
local_file_path = "my_invoice.pdf"
object_name     = "my_invoice.pdf"

# Open the file and upload it to the cloud bucket
with open(local_file_path, "rb") as f:
    object_storage_client.put_object(
        namespace,
        bucket_name,
        object_name,
        f
    )

print("✅ File uploaded successfully to OCI Object Storage!")

Output:

✅ File uploaded successfully to OCI Object Storage!

Your document is now safely sitting in Oracle Cloud, ready to be analysed! ☁️📄

📝 Step 3 — Extract All Text Using OCR

OCR stands for Optical Character Recognition. It is the ability to read text from images or scanned PDFs. Imagine taking a photo of a handwritten letter — OCR can read every single word in that photo! 📷✍️

Step 3a: Submit an OCR Job

📌 What does the code below do?

This code sends your uploaded PDF to the OCI Document Understanding AI and asks it: "Please read all the text you can find in this document." The AI reads every word on every page and saves the result back to your bucket. 📝
import oci

# Load OCI config
config = oci.config.from_file()

# Create client for Document Understanding service
ai_doc_client = oci.ai_document.AIServiceDocumentClient(config)

# Tell the AI what to do — we want TEXT EXTRACTION (OCR)
feature = oci.ai_document.models.DocumentTextExtractionFeature()

# Tell the AI where our document is stored in Object Storage
input_location = oci.ai_document.models.ObjectStorageLocations(
    object_locations=[
        oci.ai_document.models.ObjectLocation(
            namespace_name="your_namespace",       # replace with your namespace
            bucket_name="my-documents-bucket",
            object_name="my_invoice.pdf"
        )
    ]
)

# Tell the AI where to save the extracted results
output_location = oci.ai_document.models.OutputLocation(
    namespace_name="your_namespace",
    bucket_name="my-documents-bucket",
    prefix="ocr-results/"        # results go into this virtual folder
)

# Build the complete request object
create_processor_job_details = oci.ai_document.models.CreateProcessorJobDetails(
    display_name="My OCR Job",
    compartment_id="ocid1.compartment.oc1..your_compartment_id",
    input_location=input_location,
    output_location=output_location,
    processor_config=oci.ai_document.models.GeneralProcessorConfig(
        features=[feature]
    )
)

# Submit the job to OCI — this starts the AI processing
response = ai_doc_client.create_processor_job(create_processor_job_details)

print(f"✅ OCR Job started!")
print(f"   Job ID : {response.data.id}")
print(f"   Status : {response.data.lifecycle_state}")

Output:

✅ OCR Job started!
   Job ID : ocid1.aidocumentprocessorjob.oc1..examplexxxxxxx
   Status : IN_PROGRESS
✅ Good to Know — Async Processing:

Document Understanding works as an async (asynchronous) job. That means: you submit the job, OCI processes it in the background, and then you come back and collect the results when it is done. It is like dropping your clothes at a laundry shop and picking them up later! 👕

Step 3b: Wait for Completion and Read the Results

📌 What does the code below do?

This code keeps checking whether the AI has finished reading your document. Once it is done, it downloads the result file from your bucket and prints out all the text that was found — line by line, page by page! 📄
import oci
import time
import json

config = oci.config.from_file()
ai_doc_client        = oci.ai_document.AIServiceDocumentClient(config)
object_storage_client = oci.object_storage.ObjectStorageClient(config)

job_id = "ocid1.aidocumentprocessorjob.oc1..examplexxxxxxx"   # use your job ID

# Keep checking the status every 5 seconds until the job finishes
while True:
    job   = ai_doc_client.get_processor_job(job_id)
    state = job.data.lifecycle_state
    print(f"⏳ Job status: {state}")

    if state == "SUCCEEDED":
        print(" Job complete! Fetching results...")
        break
    elif state == "FAILED":
        print("❌ Job failed. Please check your inputs.")
        break

    time.sleep(5)

# Download the results JSON file from Object Storage
namespace           = "your_namespace"
bucket_name         = "my-documents-bucket"
result_object_name  = "ocr-results/my_invoice.pdf/analyzeDocumentOutput.json"

result      = object_storage_client.get_object(namespace, bucket_name, result_object_name)
result_data = json.loads(result.data.content.decode("utf-8"))

# Print the extracted text from every page
pages = result_data["pages"]
for page in pages:
    print(f"\n--- Page {page['pageNumber']} ---")
    for line in page["lines"]:
        print(line["text"])

Output (example from a real invoice):

⏳ Job status: IN_PROGRESS
⏳ Job status: IN_PROGRESS
🎉 Job complete! Fetching results...

--- Page 1 ---
INVOICE
Invoice Number: INV-2024-00123
Date: 10 April 2026
Bill To: Rohan Sharma, 42 MG Road, Jaipur
Total Amount Due: Rs. 45,500.00
Payment Due By: 30 April 2026

The AI just read your entire invoice in seconds! 🧙‍♂️

🗂️ Step 4 — Document Classification

Sometimes you receive a big pile of mixed documents — invoices, receipts, forms, ID cards — all shuffled together. Document Classification automatically sorts them! It looks at a document and says: "This is an invoice!" or "This is a passport!" 🪪

📌 What does the code below do?

This code sends your document to the AI and asks: "What type of document is this?" The AI examines it and returns a label like "INVOICE", "RECEIPT", or "DRIVING_LICENSE" — like a smart post-office worker who reads the envelope and routes every letter to the correct department! 📬
import oci

config        = oci.config.from_file()
ai_doc_client = oci.ai_document.AIServiceDocumentClient(config)

# We want DOCUMENT CLASSIFICATION this time
feature = oci.ai_document.models.DocumentClassificationFeature()

input_location = oci.ai_document.models.ObjectStorageLocations(
    object_locations=[
        oci.ai_document.models.ObjectLocation(
            namespace_name="your_namespace",
            bucket_name="my-documents-bucket",
            object_name="mystery_document.pdf"   # we do not know what this is yet!
        )
    ]
)

output_location = oci.ai_document.models.OutputLocation(
    namespace_name="your_namespace",
    bucket_name="my-documents-bucket",
    prefix="classification-results/"
)

create_processor_job_details = oci.ai_document.models.CreateProcessorJobDetails(
    display_name="Classify My Document",
    compartment_id="ocid1.compartment.oc1..your_compartment_id",
    input_location=input_location,
    output_location=output_location,
    processor_config=oci.ai_document.models.GeneralProcessorConfig(
        features=[feature]
    )
)

response = ai_doc_client.create_processor_job(create_processor_job_details)
print(f"✅ Classification job started! ID: {response.data.id}")

After the job succeeds (use the polling code from Step 3b), parse the result:

📌 What does the code below do?

This code reads the result file and prints out what document type the AI detected along with its confidence score — how sure the AI is about its answer (0% to 100%). 🎯
# After job SUCCEEDED, download and parse the result JSON
# (same polling + download pattern as Step 3b)

detected_types = result_data["detectedDocumentTypes"]
for doc_type in detected_types:
    print(f"Document Type : {doc_type['documentType']}")
    print(f"Confidence    : {doc_type['confidence'] * 100:.1f}%")

Output (example):

Document Type : INVOICE
Confidence    : 98.7%

The AI is 98.7% sure this is an invoice! That is like having a super-confident expert on your team who almost never makes a mistake. 💪

🔑 Step 5 — Key-Value Extraction

Key-Value Extraction is like filling in a form automatically. Given an invoice, it finds: Invoice Number → INV-001, Total → ₹45,000, Due Date → 15 May — all on its own! 🔑

📌 What does the code below do?

This code asks the AI to find specific named fields inside your document. Instead of reading all the text like OCR, it hunts for structured information — like a detective looking specifically for the name, amount, and date on a bill! 🕵️
import oci

config        = oci.config.from_file()
ai_doc_client = oci.ai_document.AIServiceDocumentClient(config)

# We want KEY-VALUE EXTRACTION
feature = oci.ai_document.models.DocumentKeyValueExtractionFeature()

input_location = oci.ai_document.models.ObjectStorageLocations(
    object_locations=[
        oci.ai_document.models.ObjectLocation(
            namespace_name="your_namespace",
            bucket_name="my-documents-bucket",
            object_name="my_invoice.pdf"
        )
    ]
)

output_location = oci.ai_document.models.OutputLocation(
    namespace_name="your_namespace",
    bucket_name="my-documents-bucket",
    prefix="kv-results/"
)

create_processor_job_details = oci.ai_document.models.CreateProcessorJobDetails(
    display_name="Extract Invoice Fields",
    compartment_id="ocid1.compartment.oc1..your_compartment_id",
    input_location=input_location,
    output_location=output_location,
    processor_config=oci.ai_document.models.GeneralProcessorConfig(
        features=[feature]
    )
)

response = ai_doc_client.create_processor_job(create_processor_job_details)
print(f"✅ Key-Value job started! ID: {response.data.id}")

After the job completes, parse the key-value pairs from the result:

📌 What does the code below do?

This code reads the result file and prints each detected field name, its value, and how confident the AI is — presented as a neat table. 📋
# After job SUCCEEDED and result JSON is downloaded

pages = result_data["pages"]
for page in pages:
    print(f"\n--- Page {page['pageNumber']} Key-Value Pairs ---")
    for field in page.get("documentFields", []):
        field_name  = field.get("fieldLabel",  {}).get("name",  "Unknown")
        field_value = field.get("fieldValue",  {}).get("value", "N/A")
        confidence  = field.get("fieldValue",  {}).get("confidence", 0) * 100
        print(f"  {field_name:25} → {field_value:30} ({confidence:.1f}% confident)")

Output (example):

--- Page 1 Key-Value Pairs ---
  INVOICE_NUMBER            → INV-2024-00123               (99.2% confident)
  INVOICE_DATE              → 10 April 2026                (98.5% confident)
  VENDOR_NAME               → Sharma Electronics Pvt. Ltd  (97.8% confident)
  TOTAL_AMOUNT              → Rs. 45,500.00                 (99.1% confident)
  DUE_DATE                  → 30 April 2026                (96.3% confident)

All the important fields, extracted automatically — no manual data entry needed! 

📊 Step 6 — Table Extraction

Many documents have tables inside them — like a list of items in an invoice, or a salary breakdown in a payslip. Table Extraction reads those tables and gives you clean, structured data. 📊

📌 What does the code below do?

This code reads a document and specifically targets any table it finds inside. It then prints the table data row by row — like turning a photo of a spreadsheet back into an actual spreadsheet! 🖼️ → 📊
import oci

config        = oci.config.from_file()
ai_doc_client = oci.ai_document.AIServiceDocumentClient(config)

# We want TABLE EXTRACTION — combine it with TEXT EXTRACTION
features = [
    oci.ai_document.models.DocumentTextExtractionFeature(),
    oci.ai_document.models.DocumentTableExtractionFeature()
]

input_location = oci.ai_document.models.ObjectStorageLocations(
    object_locations=[
        oci.ai_document.models.ObjectLocation(
            namespace_name="your_namespace",
            bucket_name="my-documents-bucket",
            object_name="my_invoice_with_table.pdf"
        )
    ]
)

output_location = oci.ai_document.models.OutputLocation(
    namespace_name="your_namespace",
    bucket_name="my-documents-bucket",
    prefix="table-results/"
)

create_processor_job_details = oci.ai_document.models.CreateProcessorJobDetails(
    display_name="Extract Tables from Document",
    compartment_id="ocid1.compartment.oc1..your_compartment_id",
    input_location=input_location,
    output_location=output_location,
    processor_config=oci.ai_document.models.GeneralProcessorConfig(
        features=features
    )
)

response = ai_doc_client.create_processor_job(create_processor_job_details)
print(f"✅ Table Extraction job started! ID: {response.data.id}")

After the job succeeds, read the table data:

# After job SUCCEEDED and result JSON is downloaded

pages = result_data["pages"]
for page in pages:
    tables = page.get("tables", [])
    for table_index, table in enumerate(tables):
        print(f"\n📊 Table {table_index + 1} on Page {page['pageNumber']}")
        print("-" * 60)
        for row in table["bodyRows"]:
            row_cells = [cell.get("text", "") for cell in row["cells"]]
            print("  |  ".join(f"{cell:20}" for cell in row_cells))

Output (example — invoice line items):

📊 Table 1 on Page 1
------------------------------------------------------------
  Item Name             |  Qty                 |  Unit Price          |  Amount
  Laptop Dell XPS 13    |  2                   |  Rs. 85,000          |  Rs. 170,000
  Wireless Mouse        |  5                   |  Rs. 1,200           |  Rs. 6,000
  USB-C Hub             |  3                   |  Rs. 2,500           |  Rs. 7,500

The entire product table — extracted perfectly, ready to use in your code! ✅

🏭 Step 7 — Building a Real-World Document Pipeline

Let's now combine everything we have learned into one real-world scenario. Imagine you run a small business and receive 100 invoices every week. You want to automatically read each one and save the key data into a CSV file. 📑➡️📊

📌 What does the pipeline below do?

This is a complete, end-to-end automation script that:
  1. Uploads a list of PDF invoices to OCI Object Storage
  2. Submits a batch Key-Value Extraction job to OCI Document Understanding
  3. Waits for the job to complete
  4. Downloads results and saves them into a neat CSV file
Think of it as a fully automatic invoice processing robot! 🤖
import oci
import time
import json
import csv
import os

# ──── Configuration ────────────────────────────────────────────────────────────
config             = oci.config.from_file()
ai_doc_client      = oci.ai_document.AIServiceDocumentClient(config)
os_client          = oci.object_storage.ObjectStorageClient(config)
namespace          = os_client.get_namespace().data
bucket_name        = "my-documents-bucket"
compartment_id     = "ocid1.compartment.oc1..your_compartment_id"
local_invoice_dir  = "./invoices"          # folder with your PDF invoices
output_csv         = "invoice_data.csv"


# ──── Step 1: Upload all local PDFs to Object Storage ─────────────────────────
print("📤 Uploading invoices to OCI Object Storage...")
uploaded_files = []
for filename in os.listdir(local_invoice_dir):
    if filename.endswith(".pdf"):
        with open(os.path.join(local_invoice_dir, filename), "rb") as f:
            os_client.put_object(namespace, bucket_name, filename, f)
        uploaded_files.append(filename)
        print(f"   ✅ Uploaded: {filename}")

print(f"\n📦 Total uploaded: {len(uploaded_files)} files")


# ──── Step 2: Submit batch Key-Value Extraction job ───────────────────────────
print("\n🚀 Submitting batch Key-Value Extraction job...")

object_locations = [
    oci.ai_document.models.ObjectLocation(
        namespace_name=namespace,
        bucket_name=bucket_name,
        object_name=fname
    )
    for fname in uploaded_files
]

job_details = oci.ai_document.models.CreateProcessorJobDetails(
    display_name="Batch Invoice Processor",
    compartment_id=compartment_id,
    input_location=oci.ai_document.models.ObjectStorageLocations(
        object_locations=object_locations
    ),
    output_location=oci.ai_document.models.OutputLocation(
        namespace_name=namespace,
        bucket_name=bucket_name,
        prefix="batch-results/"
    ),
    processor_config=oci.ai_document.models.GeneralProcessorConfig(
        features=[oci.ai_document.models.DocumentKeyValueExtractionFeature()]
    )
)

job      = ai_doc_client.create_processor_job(job_details)
job_id   = job.data.id
print(f"   Job ID: {job_id}")


# ──── Step 3: Poll until complete ─────────────────────────────────────────────
print("\n⏳ Waiting for job to complete...")
while True:
    state = ai_doc_client.get_processor_job(job_id).data.lifecycle_state
    print(f"   Status: {state}")
    if state in ["SUCCEEDED", "FAILED"]:
        break
    time.sleep(10)

if state == "FAILED":
    print("❌ Batch job failed.")
    exit()

print(" Job complete!")


# ──── Step 4: Download results and write CSV ──────────────────────────────────
print("\n📝 Extracting data and saving to CSV...")

with open(output_csv, "w", newline="") as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(["Filename", "Invoice Number", "Invoice Date",
                     "Vendor Name", "Total Amount", "Due Date"])

    for fname in uploaded_files:
        result_key = f"batch-results/{fname}/analyzeDocumentOutput.json"
        try:
            result_obj  = os_client.get_object(namespace, bucket_name, result_key)
            result_data = json.loads(result_obj.data.content.decode("utf-8"))
        except Exception:
            print(f"   ⚠️  Could not read result for {fname}")
            continue

        fields = {}
        for page in result_data.get("pages", []):
            for field in page.get("documentFields", []):
                key = field.get("fieldLabel", {}).get("name", "")
                val = field.get("fieldValue", {}).get("value", "")
                fields[key] = val

        writer.writerow([
            fname,
            fields.get("INVOICE_NUMBER", ""),
            fields.get("INVOICE_DATE",   ""),
            fields.get("VENDOR_NAME",    ""),
            fields.get("TOTAL_AMOUNT",   ""),
            fields.get("DUE_DATE",       "")
        ])
        print(f"   ✅ Processed: {fname}")

print(f"\n🏁 Done! Data saved to '{output_csv}'")

Output (example):

📤 Uploading invoices to OCI Object Storage...
   ✅ Uploaded: invoice_001.pdf
   ✅ Uploaded: invoice_002.pdf
   ✅ Uploaded: invoice_003.pdf

📦 Total uploaded: 3 files

🚀 Submitting batch Key-Value Extraction job...
   Job ID: ocid1.aidocumentprocessorjob.oc1..examplexxxxxxx

⏳ Waiting for job to complete...
   Status: IN_PROGRESS
   Status: IN_PROGRESS
   Status: SUCCEEDED
🎉 Job complete!

📝 Extracting data and saving to CSV...
   ✅ Processed: invoice_001.pdf
   ✅ Processed: invoice_002.pdf
   ✅ Processed: invoice_003.pdf

🏁 Done! Data saved to 'invoice_data.csv'

What used to take hours of manual data entry now takes seconds!  This is the power of OCI Document Understanding! 💪

🎓 Step 8 — Training a Custom Document Model

The built-in AI models work great for standard documents like invoices and receipts. But what if your company has its own special format — like a custom order form, a hospital discharge summary, or a legal contract? Custom Model Training lets you teach the AI to understand YOUR documents! 🧠

💡 Analogy:

Teaching a custom model is like training a new employee at your company. You show them 20–50 examples of your forms, explain what each field means, and after a few hours of training — they can read any new form perfectly! 👩‍🏫

The Training Steps (High Level)

  • Step A: Collect 20–50 sample documents of your custom type
  • Step B: Upload them to an OCI Object Storage bucket
  • Step C: Use the OCI Console (Data Labeling service) to label each field in every document
  • Step D: Create a custom model in OCI Document Understanding — it will train automatically
  • Step E: Use the trained model in your processor jobs — same code as above!
📌 What does the code below do?

This code creates a custom model training job in OCI. It tells OCI: "Here is a folder of labelled sample documents. Please learn from them and build a model I can use for future documents." 🏋️
import oci

config        = oci.config.from_file()
ai_doc_client = oci.ai_document.AIServiceDocumentClient(config)

# Step 1: First create a Project (a container for your models)
project_response = ai_doc_client.create_project(
    oci.ai_document.models.CreateProjectDetails(
        display_name="My Invoice Recognizer Project",
        compartment_id="ocid1.compartment.oc1..your_compartment_id"
    )
)
project_id = project_response.data.id
print(f"✅ Project created: {project_id}")


# Step 2: Create and start the model training job
model_response = ai_doc_client.create_model(
    oci.ai_document.models.CreateModelDetails(
        display_name="Custom Invoice Model v1",
        compartment_id="ocid1.compartment.oc1..your_compartment_id",
        project_id=project_id,
        model_type="KEY_VALUE_EXTRACTION",     # or "DOCUMENT_CLASSIFICATION"
        training_dataset=oci.ai_document.models.DatasetSummary(
            namespace_name="your_namespace",
            bucket_name="my-training-data-bucket",
            prefix="labelled-invoices/"         # folder with labelled sample docs
        )
    )
)

model_id = model_response.data.id
print(f"🚀 Model training started! Model ID: {model_id}")
print(f"   Status: {model_response.data.lifecycle_state}")
print(f"\n⏳ Training may take 30–60 minutes. Grab a coffee! ☕")

Output:

✅ Project created: ocid1.aidocumentproject.oc1..examplexxxxxxx
🚀 Model training started! Model ID: ocid1.aidocumentmodel.oc1..examplexxxxxxx
   Status: CREATING

⏳ Training may take 30–60 minutes. Grab a coffee! ☕
✅ Once the model is ACTIVE:

Simply pass your model_id into any processor job using the model_id parameter — and the AI will use your custom-trained model instead of the default one! 🎯

🏆 Best Practices

  • 📷 Always use high-quality scans — minimum 200 DPI for best OCR accuracy. Blurry photos give blurry results!
  • 🗂️ Combine multiple features in one job — you can request Text + Table + Key-Value extraction all in a single API call. No need for three separate jobs!
  • 🔄 Use batch processing — submit multiple documents in one job instead of submitting one at a time. It is faster and cheaper!
  • 🧪 Test with sample documents first — before running on production data, always test with a handful of examples to verify accuracy.
  • 💾 Store results in a database — after extraction, save your structured data to Autonomous Database or MySQL HeatWave for easy querying and dashboards.
  • 🔐 Use OCI Vault for secrets — never hardcode your compartment ID or API keys directly in your code. Use environment variables or OCI Vault.
❌ Common Mistakes to Avoid:

  • Do NOT submit a job and immediately try to read results — the job runs asynchronously, so always poll for completion first!
  • Do NOT use very small or compressed images — OCR accuracy drops sharply below 150 DPI
  • Do NOT ignore the confidence scores — always set a minimum threshold (e.g., 80%) before trusting extracted values

🌍 Real-World Use Cases

  • 🏦 Banking & Finance: Auto-process loan applications, KYC documents, and cheques
  • 🏥 Healthcare: Extract patient info from prescription pads and discharge summaries
  • 🛒 Retail & E-Commerce: Process supplier invoices and delivery receipts automatically
  • 🏛️ Government: Digitise old paper records and process citizen applications faster
  • 📦 Logistics: Read shipping labels and customs documents in real time
  • ⚖️ Legal: Extract clauses and dates from long contracts without reading every word

📝 Quick Summary — What We Learned

  • What OCI Document Understanding is → An AI service that reads and understands any document
  • OCR (Text Extraction) → Reads all visible text from any PDF or image
  • Document Classification → Identifies what type of document you have
  • Key-Value Extraction → Pulls out named fields like Invoice Number, Total Amount, Due Date
  • Table Extraction → Reads structured tables inside documents
  • Custom Model Training → Teach the AI to understand your own document formats
  • Batch Pipeline → Process hundreds of documents automatically using Python

Happy building! 📄✨

Comments