Skip to main content

Invoice Compliance & Fraud Detection: Build an AI-Powered Capstone Project

Calculating read time…

Today we are going to build something really cool — a system that can READ text from images (like invoices and documents) and turn it into neat, organized data automatically!

We will build this step by step, like building with LEGO bricks 🧱. Each section adds one more piece until we have a complete, working AI service!

💡 What we are building: You give it a photo of an invoice → Our system reads it → It gives back clean, structured data in JSON format. Like magic! ✨



🗺️ Project Overview — The Big Picture

Before we start coding, let's understand what we are building. Think of it like a factory 🏭:

┌─────────────────────────────────────────────────────────────┐
│              OCI Vision OCR Inference Service               │
│                                                             │
│  📷 Image (Invoice)                                         │
│       │                                                     │
│       ▼                                                     │
│  🌐 FastAPI (api.py)  ← You send image here via HTTP       │
│       │                                                     │
│       ▼                                                     │
│  🧠 score.py  ← The BRAIN — processes everything           │
│       │                                                     │
│       ▼                                                     │
│  ☁️  OCI Vision AI  ← Oracle's smart robot reads image     │
│       │                                                     │
│       ▼                                                     │
│  📋 JSON Output  ← Clean, structured data returned         │
└─────────────────────────────────────────────────────────────┘

📁 Files We Will Create

  • score.py → The brain of our system (core logic)
  • api.py → The front door (FastAPI web server)
  • requirements.txt → Shopping list of tools
  • Dockerfile → Packing everything into a box
  • .env → Secret settings file

🔹 SECTION 1 — What is score.py? (The Brain! 🧠)

Imagine you have a very smart robot assistant. You give the robot a photo of your homework. The robot reads it, thinks about it, and gives you back the answers in a neat list.

score.py is exactly that robot's brain.

In the Machine Learning world, score.py is a special file that:

  • 📥 Receives the input (our invoice image)
  • 🔄 Processes it (calls OCI Vision to read the text)
  • 📤 Returns the result (clean JSON data)

Think of score.py like the recipe book 📖 in a restaurant kitchen. Every time a customer orders food (sends an image), the chef (our code) follows the recipe (score.py) to make the dish (JSON output).

⭐ Why is it called "score"? Because in ML, "scoring" means running a prediction or analysis on new data. It's like a teacher "scoring" your exam paper!


🔹 SECTION 2 — What is OCI Vision? (The Smart Reading Robot 📖)

OCI Vision is Oracle Cloud's Artificial Intelligence service. Think of it as a super-smart robot 🤖 that has been trained to look at millions of images and learned to read text from them.

🧒 Simple explanation: You know how you can look at a handwritten note and read what it says? OCI Vision can do the same thing — but for computers! It looks at an image (like an invoice photo) and reads all the text it finds.

What can OCI Vision do?

  • 📝 OCR (Text Reading) → Reads ALL text from an image
  • 🔑 Key-Value Extraction → Finds pairs like "Invoice No: 12345" or "Total: $500"
  • 📊 Table Extraction → Reads tables and keeps rows and columns organized
  • 📄 Document Classification → Recognizes if the document is an invoice, receipt, etc.

How does OCR work? (Super simple!)

🖼️  You give it a photo:
    ┌────────────────────┐
    │  INVOICE #1234     │
    │  Date: 2026-04-01  │
    │  Total: $500.00    │
    └────────────────────┘

🤖  OCI Vision reads it and gives back:
    "INVOICE #1234 Date: 2026-04-01 Total: $500.00"

🧹  Our score.py then organizes it:
    {
      "invoice_number": "1234",
      "date": "2026-04-01",
      "total": "$500.00"
    }

It's like having a super-fast, super-accurate typist who can read any document and type it out for you! ⌨️


🛠️ SECTION 2.5 — Setup Before Coding

Before we write any code, we need to set up some things in Oracle Cloud. Think of it like getting your tools ready before building something! 🔧

Step A — Get Your OCI Credentials

OCI Vision needs to know WHO is asking it to read documents. It uses a special key (like a house key 🔑) to verify it's really you.

  1. Log in to OCI Console at cloud.oracle.com
  2. Click your Profile icon (top right) → My Profile
  3. Go to API Keys → Click "Add API Key"
  4. Download the Private Key file (save it safely!)
  5. Copy the config snippet Oracle shows you — it looks like this:
[DEFAULT]
user=ocid1.user.oc1..aaaaaaaXXXXXXXXXXXXXXXX
fingerprint=aa:bb:cc:dd:ee:ff:11:22:33:44
tenancy=ocid1.tenancy.oc1..aaaaaaaXXXXXXXXXXXXXXXX
region=us-ashburn-1
key_file=~/.oci/oci_api_key.pem

Step B — Create a .env File (Secret Settings)

We store our secrets in a .env file so we don't write them in the code. Think of it like a private diary 📓 that only our app can read!

📋 What this .env file does:
It stores our secret settings (like passwords and IDs) in one safe place. Our code reads from here instead of having secrets hardcoded inside the code itself.

# .env — Secret settings file (NEVER share this file publicly!)

# ─── OCI Authentication ────────────────────────────────────
OCI_USER_OCID=ocid1.user.oc1..aaaaaaaXXXXXXXXXXXXX
OCI_FINGERPRINT=aa:bb:cc:dd:ee:ff:11:22:33:44
OCI_TENANCY_OCID=ocid1.tenancy.oc1..aaaaaaaXXXXXXXXXXXXX
OCI_REGION=us-ashburn-1
OCI_KEY_FILE=./oci_api_key.pem

# ─── OCI Compartment ───────────────────────────────────────
# This is like a folder in OCI where your resources live
COMPARTMENT_ID=ocid1.compartment.oc1..aaaaaaaXXXXXXXXXXXXX

# ─── Service Settings ──────────────────────────────────────
# Port our API will run on
APP_PORT=8000

⚠️ Important: Add .env to your .gitignore file! Never push this to GitHub. It's like your house key — don't share it!


🔹 SECTION 3 — Building score.py (The Brain! 🧠)

Now the real fun begins! We are going to build the brain of our system. This is the most important file — everything else talks to score.py.

We will build score.py in 4 parts, one at a time. Like building a sandwich — one layer at a time! 🥪

Part 1 of score.py — Imports and Setup

📋 Step 1 — What this code will do:
This is the very top of our file. We are importing (bringing in) all the tools we need. Think of it like opening your school bag and taking out your pencil, ruler, and eraser before you start homework!

📋 Step 2 — Why we need it:
Without importing these tools, our code has no power. Python itself is basic — these imports give it superpowers like talking to OCI Vision, reading files, and working with JSON.

# ═══════════════════════════════════════════════════════════════
# score.py — The Brain of our OCI Vision OCR Inference Service
# ═══════════════════════════════════════════════════════════════

# ─── Standard Python libraries (already included in Python) ───
import os          # To read environment variables (our settings)
import base64      # To convert images into text format for sending
import json        # To work with JSON data (our output format)
import logging     # To print helpful messages so we know what's happening
import re          # To do pattern matching (find things like dates, amounts)

# ─── OCI (Oracle Cloud) Python SDK ────────────────────────────
# This is our connection to Oracle Cloud services
import oci
import oci.ai_vision          # The OCI Vision AI service
import oci.ai_document        # The OCI Document Understanding service

# ─── For loading our .env settings file ───────────────────────
from dotenv import load_dotenv   # Reads our .env file

# ─── Type hints (helps us write cleaner code) ─────────────────
from typing import Dict, Any, Optional

# ─── Load all settings from our .env file ─────────────────────
load_dotenv()

# ─── Set up logging so we can see what's happening ────────────
# This is like having a "diary" that writes down every step
logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s [%(levelname)s] %(message)s"
)
logger = logging.getLogger(__name__)

📋 Step 4 — Explaining important lines:

  • import oci → This brings in the entire Oracle Cloud SDK — our toolkit for talking to OCI
  • import base64 → Images need to be converted to text (Base64 format) before sending to OCI Vision. Think of it like translating a photo into a secret code!
  • load_dotenv() → This reads our .env file and loads all our secret settings
  • logging.basicConfig() → This sets up our diary so we can see messages like "Image received!" or "Error happened here!"

Part 2 of score.py — Configuration Class

📋 Step 1 — What this code will do:
We create a class (a blueprint 📐) that holds all our OCI connection settings. Think of it like filling out a form with your name, address, and phone number — but this form is for connecting to Oracle Cloud!

📋 Step 2 — Why we need it:
OCI Vision needs to know WHO is calling it (our credentials) before it does anything. This class bundles all those details neatly in one place.

# ═══════════════════════════════════════════════════════════════
# PART 2: Configuration — Holds all our OCI settings
# ═══════════════════════════════════════════════════════════════

class OCIConfig:
    """
    This class stores all the settings needed to connect to OCI.
    Think of it like a contact card for Oracle Cloud!
    """

    def __init__(self):
        # ── Read all settings from .env file ──────────────────
        # os.getenv() reads a value from our .env file
        # The second argument is the "default" if the variable is missing
        self.user        = os.getenv("OCI_USER_OCID", "")
        self.fingerprint = os.getenv("OCI_FINGERPRINT", "")
        self.tenancy     = os.getenv("OCI_TENANCY_OCID", "")
        self.region      = os.getenv("OCI_REGION", "us-ashburn-1")
        self.key_file    = os.getenv("OCI_KEY_FILE", "./oci_api_key.pem")

        # ── The compartment is like a folder in OCI ────────────
        self.compartment_id = os.getenv("COMPARTMENT_ID", "")

        # ── Validate that required settings are present ────────
        self._validate()

    def _validate(self):
        """Check that all required settings are filled in."""
        missing = []

        # Check each required setting — if empty, add to "missing" list
        if not self.user:           missing.append("OCI_USER_OCID")
        if not self.fingerprint:    missing.append("OCI_FINGERPRINT")
        if not self.tenancy:        missing.append("OCI_TENANCY_OCID")
        if not self.compartment_id: missing.append("COMPARTMENT_ID")

        # If anything is missing, raise an error (stop the program)
        if missing:
            raise ValueError(
                f"❌ Missing required settings in .env file: {', '.join(missing)}\n"
                f"   Please fill them in and try again!"
            )

        logger.info("✅ OCI Configuration loaded successfully!")

    def to_oci_dict(self) -> dict:
        """
        Convert our settings into the format OCI SDK expects.
        This is like translating our settings into OCI's language!
        """
        return {
            "user":        self.user,
            "fingerprint": self.fingerprint,
            "tenancy":     self.tenancy,
            "region":      self.region,
            "key_file":    self.key_file,
        }

📋 Step 4 — Explaining important lines:

  • os.getenv("OCI_USER_OCID", "") → Read the value of OCI_USER_OCID from our .env file. If it's not there, use empty string ""
  • self._validate() → This runs automatically when we create the config — like a checklist before takeoff ✈️
  • def to_oci_dict() → Converts our settings into a Python dictionary that OCI SDK understands

Part 3 of score.py — The OCR Engine (The Real Magic! ✨)

📋 Step 1 — What this code will do:
This is the main worker! It takes an image, sends it to OCI Vision, gets back the text, and organizes everything into clean JSON. This is like the kitchen chef 👨‍🍳 who takes the ingredients (image) and creates the final dish (JSON)!

📋 Step 2 — Why we need it:
This is the core "inference" logic — the actual intelligence of our service. Without this, we're just accepting images but doing nothing with them!

# ═══════════════════════════════════════════════════════════════
# PART 3: OCREngine — The main worker class that does all the work
# ═══════════════════════════════════════════════════════════════

class OCREngine:
    """
    This is the main engine of our service.
    It connects to OCI Vision and extracts text from images.

    Think of it like a smart photocopier that not only copies
    the document but also understands and organizes what it reads!
    """

    def __init__(self, config: OCIConfig):
        """Set up the engine with OCI connection."""
        self.config = config

        # ── Connect to OCI Vision service ──────────────────────
        # This is like calling Oracle Cloud on the phone!
        oci_config = config.to_oci_dict()

        # Create the OCI Vision client (our connection to AI)
        self.vision_client = oci.ai_vision.AIServiceVisionClient(
            config=oci_config
        )

        # Create the OCI Document Understanding client
        # (better for structured documents like invoices!)
        self.doc_client = oci.ai_document.AIServiceDocumentClient(
            config=oci_config
        )

        logger.info("✅ OCR Engine initialized — connected to OCI Vision!")

    def image_to_base64(self, image_path: str) -> str:
        """
        Convert an image file into Base64 text format.

        Why Base64? Imagine you want to send a photo via text message.
        You can't send a photo as raw bytes in JSON, so we encode it
        into a long string of letters and numbers first!

        It's like translating a photo into a secret code 🔐
        """
        logger.info(f"📷 Reading image from: {image_path}")

        # Open the image file in binary mode (rb = read binary)
        with open(image_path, "rb") as image_file:
            # Read all the bytes of the image
            image_bytes = image_file.read()

        # Convert those bytes into a Base64 string
        base64_string = base64.b64encode(image_bytes).decode("utf-8")

        logger.info(f"✅ Image converted to Base64 ({len(base64_string)} chars)")
        return base64_string

    def image_bytes_to_base64(self, image_bytes: bytes) -> str:
        """
        Same as above but works with bytes directly (used by our API).
        This is when the image is uploaded directly via HTTP request.
        """
        return base64.b64encode(image_bytes).decode("utf-8")

    def call_oci_vision_ocr(self, base64_image: str) -> dict:
        """
        Call OCI Vision API to extract text from the image.

        This is the most important step — we send the image to
        Oracle's AI brain and get back all the text it found!

        Think of it like dropping your photo at a copy shop 📸
        and they type out everything written in the photo for you!
        """
        logger.info("🤖 Sending image to OCI Vision for OCR analysis...")

        # ── Build the request for OCI Vision ──────────────────
        # We tell OCI Vision: "Here is the image (as Base64),
        # please extract the text from it!"

        analyze_details = oci.ai_vision.models.AnalyzeDocumentDetails(

            # ── The image data ─────────────────────────────────
            # "INLINE" means we send the image directly (not from storage)
            document=oci.ai_vision.models.InlineDocumentDetails(
                source="INLINE",
                data=base64_image    # Our Base64 encoded image
            ),

            # ── What we want OCI to do ─────────────────────────
            # We request THREE types of analysis at once!
            features=[
                # 1. READ ALL TEXT from the document
                oci.ai_vision.models.DocumentTextExtractionFeature(
                    feature_type="TEXT_DETECTION"
                ),
                # 2. FIND KEY-VALUE PAIRS (like "Invoice No: 1234")
                oci.ai_vision.models.DocumentKeyValueExtractionFeature(
                    feature_type="KEY_VALUE_DETECTION"
                ),
                # 3. CLASSIFY what type of document it is
                oci.ai_vision.models.DocumentClassificationFeature(
                    feature_type="DOCUMENT_CLASSIFICATION"
                ),
            ],

            # Our compartment ID (where our OCI resources live)
            compartment_id=self.config.compartment_id,
        )

        # ── Make the actual API call to OCI Vision ─────────────
        # This is like pressing the "send" button!
        response = self.vision_client.analyze_document(
            analyze_document_details=analyze_details
        )

        logger.info("✅ OCI Vision responded successfully!")
        return response.data

    def extract_full_text(self, vision_response) -> str:
        """
        Extract ALL the text from OCI Vision's response.

        OCI Vision returns text organized by pages and blocks.
        We join it all together into one big string.

        Think of it like collecting puzzle pieces and joining them! 🧩
        """
        all_text_parts = []

        # Loop through each page in the document
        if hasattr(vision_response, 'pages') and vision_response.pages:
            for page in vision_response.pages:
                # Each page has multiple "lines" of text
                if hasattr(page, 'lines') and page.lines:
                    for line in page.lines:
                        if hasattr(line, 'text') and line.text:
                            all_text_parts.append(line.text)

        # Join all pieces together with newlines between them
        full_text = "\n".join(all_text_parts)
        return full_text

    def extract_key_values(self, vision_response) -> Dict[str, Any]:
        """
        Extract KEY-VALUE PAIRS from OCI Vision's response.

        OCI Vision automatically detects fields like:
        - "Invoice Number" → "INV-2026-001"
        - "Date" → "2026-04-18"
        - "Total Amount" → "$1,500.00"

        Think of it like a form where you have label → value pairs!
        📝 Name: John   →   key="Name", value="John"
        """
        key_values = {}

        # Check if OCI found any key-value pairs
        if hasattr(vision_response, 'detected_document_types'):
            logger.info(
                f"📄 Document type: {vision_response.detected_document_types}"
            )

        # Extract key-value pairs if OCI found them
        if hasattr(vision_response, 'key_value_fields') and vision_response.key_value_fields:
            logger.info(f"🔑 Found {len(vision_response.key_value_fields)} key-value fields")

            for field in vision_response.key_value_fields:
                # Get the field name (key) and its value
                key   = field.name.lower().replace(" ", "_") if field.name else "unknown"
                value = field.value if hasattr(field, 'value') and field.value else ""
                confidence = field.confidence if hasattr(field, 'confidence') else 0

                # Only include results with reasonable confidence (> 50%)
                if confidence > 0.5:
                    key_values[key] = {
                        "value":      value,
                        "confidence": round(confidence * 100, 1)  # Convert to percentage
                    }
        else:
            logger.info("ℹ️  No pre-built key-value fields found. Will use pattern matching.")

        return key_values

    def extract_with_patterns(self, full_text: str) -> Dict[str, Any]:
        """
        Use PATTERN MATCHING to find common invoice fields.

        Sometimes OCI Vision reads the text but doesn't automatically
        label the fields. In that case, we use REGEX patterns to
        find common patterns in the text ourselves.

        Think of it like being a detective 🕵️ looking for clues!
        "Anywhere I see 'Invoice #' followed by numbers — that's the invoice number!"
        """
        extracted = {}

        # ── Pattern 1: Invoice Number ──────────────────────────
        # Looks for: "Invoice No: 1234" or "Invoice #1234" or "Inv: 1234"
        invoice_pattern = r'(?:invoice\s*(?:no|number|#|num)?[:.\s#]*)\s*([A-Z0-9\-]+)'
        match = re.search(invoice_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["invoice_number"] = match.group(1).strip()

        # ── Pattern 2: Date ────────────────────────────────────
        # Looks for dates like: 2026-04-18 or 04/18/2026 or 18-Apr-2026
        date_pattern = r'(?:date|dated|invoice date)?[:.\s]*(\d{1,2}[\/\-\.]\d{1,2}[\/\-\.]\d{2,4}|\d{4}[\/\-\.]\d{2}[\/\-\.]\d{2})'
        match = re.search(date_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["date"] = match.group(1).strip()

        # ── Pattern 3: Total Amount ────────────────────────────
        # Looks for: "Total: $500.00" or "Grand Total: 1,500"
        total_pattern = r'(?:grand\s*total|total\s*amount|total|amount\s*due)[:.\s]*[$₹€£]?\s*([\d,]+\.?\d{0,2})'
        match = re.search(total_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["total_amount"] = match.group(1).strip()

        # ── Pattern 4: Tax Amount ──────────────────────────────
        tax_pattern = r'(?:tax|gst|vat|cgst|sgst|igst)[:.\s]*[$₹€£]?\s*([\d,]+\.?\d{0,2})'
        match = re.search(tax_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["tax_amount"] = match.group(1).strip()

        # ── Pattern 5: Vendor / From ───────────────────────────
        vendor_pattern = r'(?:from|vendor|seller|company|bill\s*from)[:.\s]+([A-Za-z0-9\s\.\,\&]+?)(?:\n|$)'
        match = re.search(vendor_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["vendor_name"] = match.group(1).strip()

        # ── Pattern 6: Email Address ───────────────────────────
        email_pattern = r'[\w\.-]+@[\w\.-]+\.\w{2,}'
        match = re.search(email_pattern, full_text)
        if match:
            extracted["email"] = match.group(0)

        # ── Pattern 7: Phone Number ────────────────────────────
        phone_pattern = r'(?:phone|tel|mob|contact)?[:.\s]*(\+?[\d\s\-\(\)]{10,15})'
        match = re.search(phone_pattern, full_text, re.IGNORECASE)
        if match:
            extracted["phone"] = match.group(1).strip()

        return extracted

📋 Step 4 — Explaining important lines:

  • source="INLINE" → We send the image directly inside the API request (no need for separate cloud storage)
  • feature_type="TEXT_DETECTION" → Tells OCI Vision: "Please read ALL the text in this image"
  • feature_type="KEY_VALUE_DETECTION" → Tells OCI Vision: "Please find field-value pairs like Invoice Number, Date, Total"
  • confidence > 0.5 → Only keep results OCI Vision is more than 50% sure about
  • re.search(pattern, text) → The detective 🕵️ looking for patterns in the text using regex

Part 4 of score.py — The Main Score Function

📋 Step 1 — What this code will do:
This is the final piece! The score() function is what gets called when someone sends us an image. It orchestrates everything — calling the right functions in the right order and returning the final JSON result.

📋 Step 2 — Why we need it:
This is the "main entrance" to our service. Think of it like the front desk 🎪 of a hotel — all guests (requests) come here first, and the front desk directs them to the right rooms (functions).

# ═══════════════════════════════════════════════════════════════
# PART 4: The Main score() function — Entry point of our service
# ═══════════════════════════════════════════════════════════════

def score(image_input: Any, input_type: str = "path") -> Dict[str, Any]:
    """
    THE MAIN FUNCTION — This is what gets called to process an image.

    This is like the "Play" button ▶️ of our service.
    Press it with an image → get back JSON data.

    Arguments:
        image_input: Either a file path (str) or raw image bytes
        input_type:  "path" if you give a file path
                     "bytes" if you give raw image bytes

    Returns:
        A dictionary (JSON) with all extracted information
    """
    logger.info("=" * 60)
    logger.info("🚀 score() called — Starting OCR inference...")
    logger.info("=" * 60)

    # ── Step 1: Initialize configuration ──────────────────────
    # Load our OCI settings from .env file
    try:
        config = OCIConfig()
    except ValueError as e:
        logger.error(f"❌ Configuration error: {e}")
        return {
            "success": False,
            "error":   str(e),
            "data":    {}
        }

    # ── Step 2: Create the OCR Engine ─────────────────────────
    try:
        engine = OCREngine(config)
    except Exception as e:
        logger.error(f"❌ Failed to initialize OCR Engine: {e}")
        return {
            "success": False,
            "error":   f"Engine initialization failed: {str(e)}",
            "data":    {}
        }

    # ── Step 3: Convert image to Base64 ───────────────────────
    # OCI Vision needs the image in Base64 format
    try:
        if input_type == "path":
            # Image given as a file path (like "/home/user/invoice.jpg")
            if not os.path.exists(image_input):
                raise FileNotFoundError(f"Image file not found: {image_input}")
            base64_image = engine.image_to_base64(image_input)

        elif input_type == "bytes":
            # Image given as raw bytes (from HTTP upload)
            base64_image = engine.image_bytes_to_base64(image_input)

        else:
            raise ValueError(f"Unknown input_type: {input_type}. Use 'path' or 'bytes'")

    except Exception as e:
        logger.error(f"❌ Image preparation failed: {e}")
        return {
            "success": False,
            "error":   f"Image error: {str(e)}",
            "data":    {}
        }

    # ── Step 4: Call OCI Vision OCR ───────────────────────────
    # Send image to Oracle's AI and get text back
    try:
        vision_response = engine.call_oci_vision_ocr(base64_image)
    except Exception as e:
        logger.error(f"❌ OCI Vision API call failed: {e}")
        return {
            "success": False,
            "error":   f"OCI Vision error: {str(e)}",
            "data":    {}
        }

    # ── Step 5: Extract all text ───────────────────────────────
    full_text = engine.extract_full_text(vision_response)
    logger.info(f"📝 Extracted {len(full_text)} characters of text")

    # ── Step 6: Extract key-value pairs ───────────────────────
    # First try OCI's built-in key-value detection
    oci_key_values = engine.extract_key_values(vision_response)

    # Then use our pattern matching as a backup / supplement
    pattern_key_values = engine.extract_with_patterns(full_text)

    # ── Step 7: Get document type ─────────────────────────────
    doc_type = "unknown"
    doc_confidence = 0
    if hasattr(vision_response, 'detected_document_types') and vision_response.detected_document_types:
        doc_type = vision_response.detected_document_types[0].document_type
        doc_confidence = round(
            vision_response.detected_document_types[0].confidence * 100, 1
        )

    # ── Step 8: Build the final JSON response ─────────────────
    # This is our "finished dish" 🍽️ ready to serve!
    result = {
        "success": True,
        "document": {
            "type":            doc_type,
            "type_confidence": f"{doc_confidence}%",
            "total_characters": len(full_text),
            "total_lines":      full_text.count('\n') + 1,
        },
        "extracted_text": {
            "full_text":   full_text,
            "line_count":  full_text.count('\n') + 1,
        },
        "key_values": {
            # From OCI Vision's built-in detection
            "oci_detected":      oci_key_values,
            # From our custom pattern matching
            "pattern_extracted": pattern_key_values,
        },
        "summary": {
            # Merge both sources — OCI values take priority
            **pattern_key_values,
            **{k: v["value"] for k, v in oci_key_values.items() if v.get("value")}
        }
    }

    logger.info("✅ score() completed successfully!")
    logger.info(f"   Document type: {doc_type} ({doc_confidence}% confidence)")
    logger.info(f"   Text length: {len(full_text)} chars")
    logger.info(f"   Key-values found: {len(oci_key_values)} OCI + {len(pattern_key_values)} pattern")

    return result


# ─── Allow running score.py directly for testing ──────────────
if __name__ == "__main__":
    # This block only runs when you do: python score.py
    # Great for quick testing! 🧪
    import sys

    if len(sys.argv) > 1:
        # If you provide an image path as argument: python score.py myinvoice.jpg
        image_path = sys.argv[1]
        print(f"\n🧪 Testing score.py with image: {image_path}\n")
        result = score(image_path, input_type="path")
    else:
        print("Usage: python score.py ")
        print("Example: python score.py invoice.jpg")
        sys.exit(1)

    # Print the result as pretty JSON
    print("\n" + "=" * 60)
    print("📋 RESULT:")
    print("=" * 60)
    print(json.dumps(result, indent=2))

📋 Step 4 — Explaining important lines:

  • def score(image_input, input_type="path") → Our main function that accepts both file paths and raw bytes
  • try/except blocks → Like safety nets 🥅. If anything goes wrong at any step, we catch the error and return a nice error message instead of crashing!
  • **pattern_key_values, **{k: v["value"]...} → Merges two dictionaries together. OCI Vision results override pattern results (OCI is more reliable!)
  • if __name__ == "__main__" → This makes score.py also runnable directly from the terminal for testing

🔹 SECTION 4 — requirements.txt (The Shopping List 🛒)

📋 Step 1 — What this file does:
requirements.txt is like a shopping list 🛒. When you want to cook a recipe, you first write down all the ingredients you need. Then you go to the shop and buy them.

Similarly, requirements.txt lists all the Python libraries (tools) our app needs. When someone wants to run our app, they just run one command and Python downloads everything automatically!

📋 Step 2 — Why we need it:
Without this file, someone copying our code would have to guess which libraries to install. With this file, they just run pip install -r requirements.txt and everything is ready!

# requirements.txt
# ─────────────────────────────────────────────────────────────
# Shopping list of ALL tools our OCI Vision OCR service needs
# Install all at once with: pip install -r requirements.txt
# ─────────────────────────────────────────────────────────────

# ── Oracle Cloud Infrastructure SDK ───────────────────────────
# This is our main connection to Oracle Cloud services
# Includes OCI Vision, Document Understanding, and more
oci>=2.120.0

# ── FastAPI — Our Web Server Framework ────────────────────────
# This turns our Python code into a real HTTP API server
# Think of it like the "front door" of our service
fastapi>=0.110.0

# ── Uvicorn — FastAPI needs this to actually run ──────────────
# It's the engine that makes FastAPI run
# Like the engine of a car — FastAPI is the car body!
uvicorn[standard]>=0.29.0

# ── Python-multipart — For handling file uploads ──────────────
# When someone uploads an image via HTTP, this handles it
python-multipart>=0.0.9

# ── Python-dotenv — For reading our .env file ─────────────────
# Reads secret settings from the .env file
python-dotenv>=1.0.0

# ── Pillow — For image processing ─────────────────────────────
# Lets us open, resize, and check images
# Like having image editing tools built in
Pillow>=10.3.0

# ── Pydantic — For data validation ────────────────────────────
# Makes sure the data we send and receive is in the right format
# Like a quality checker for our data
pydantic>=2.6.0

# ── Requests — For making HTTP calls (optional) ───────────────
# Useful for testing our API
requests>=2.31.0

# ── Pytest — For running tests ────────────────────────────────
pytest>=8.0.0
pytest-asyncio>=0.23.0

📋 Important libraries explained:

  • oci → The Oracle Cloud SDK — this is what lets us talk to OCI Vision
  • fastapi → Creates our web API — turns Python functions into HTTP endpoints
  • uvicorn → The web server that actually runs FastAPI (like Apache/Nginx but simpler)
  • python-multipart → Handles file uploads (when someone uploads an image via HTTP)
  • python-dotenv → Reads our .env file containing secrets
  • Pillow → Python image library — helps us validate and process images

💡 Install command:

# Run this command in your terminal to install everything:
pip install -r requirements.txt

# You'll see lots of packages being downloaded and installed.
# When it says "Successfully installed..." — you're ready! ✅

🔹 SECTION 5 — API Wrapper with FastAPI (The Front Door 🚪)

📋 Step 1 — What this code will do:
Now we build the API wrapper — a web server that exposes our score.py as an HTTP service. Think of it like building the reception desk 🎪 of a hotel. Guests (clients) come to the reception (API) and it handles their requests!

📋 Step 2 — Why we need it:
score.py is the brain, but it can't receive HTTP requests by itself. FastAPI creates a web server with endpoints (URLs) that anyone can call with an HTTP request to use our OCR service.

The main endpoint we expose:

  • POST /extract → Upload an image file, get back JSON with extracted text
  • GET /health → Check if the service is alive and working
  • GET / → Home page with API documentation
# ═══════════════════════════════════════════════════════════════
# api.py — The Front Door (FastAPI Web Server)
# ═══════════════════════════════════════════════════════════════
# This file creates our HTTP API so anyone can use our OCR service
# by sending a simple web request — just like using a website!

# ─── Import FastAPI and related tools ─────────────────────────
from fastapi import FastAPI, File, UploadFile, HTTPException, Request
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware
import uvicorn
import logging
import time
import os
from typing import Optional
from pydantic import BaseModel

# ─── Import our score.py (the brain!) ─────────────────────────
from score import score

# ─── Load environment settings ────────────────────────────────
from dotenv import load_dotenv
load_dotenv()

# ─── Set up logging ───────────────────────────────────────────
logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s [%(levelname)s] %(message)s"
)
logger = logging.getLogger(__name__)

# ═══════════════════════════════════════════════════════════════
# Create the FastAPI application
# This is like opening the doors of our shop!
# ═══════════════════════════════════════════════════════════════
app = FastAPI(
    title="OCI Vision OCR Inference Service",
    description="""
    🤖 An ML-style inference service using OCI Vision OCR.

    This service accepts invoice/document images and returns
    structured JSON with extracted text and key-value pairs.

    ## How to use:
    1. POST an image to `/extract`
    2. Get back structured JSON data

    ## Supported formats:
    JPEG, PNG, TIFF, PDF (first page only)
    """,
    version="1.0.0",
    docs_url="/docs",       # Swagger UI at /docs
    redoc_url="/redoc",     # ReDoc at /redoc
)

# ─── Allow requests from any origin (CORS) ────────────────────
# This lets web browsers and other services call our API
# Think of it like allowing visitors from all countries 🌍
app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],      # Allow all origins
    allow_credentials=True,
    allow_methods=["*"],      # Allow all HTTP methods (GET, POST, etc.)
    allow_headers=["*"],      # Allow all headers
)

# ═══════════════════════════════════════════════════════════════
# ROUTE 1: Health Check — Is our service alive?
# Like checking if the shop is open! 🟢
# ═══════════════════════════════════════════════════════════════
@app.get(
    "/health",
    summary="Health Check",
    description="Check if the service is running correctly"
)
async def health_check():
    """
    A simple health check endpoint.

    Returns a JSON response showing the service status.
    Useful for monitoring, Docker health checks, Kubernetes probes.
    """
    return {
        "status":    "healthy",
        "service":   "OCI Vision OCR Inference",
        "version":   "1.0.0",
        "timestamp": time.time(),
        "message":   "🤖 Service is running! Ready to read your documents."
    }


# ═══════════════════════════════════════════════════════════════
# ROUTE 2: Home Page
# ═══════════════════════════════════════════════════════════════
@app.get(
    "/",
    summary="Home",
    description="Welcome page with service information"
)
async def home():
    """Welcome page — shows API information."""
    return {
        "service":     "OCI Vision OCR Inference Service",
        "version":     "1.0.0",
        "description": "Upload invoice/document images and extract structured data",
        "endpoints": {
            "POST /extract":   "Upload an image → get structured JSON",
            "GET  /health":    "Check service health",
            "GET  /docs":      "Interactive API documentation (Swagger)",
            "GET  /redoc":     "Alternative API documentation",
        },
        "usage_example": {
            "command": "curl -X POST http://localhost:8000/extract -F 'file=@invoice.jpg'",
            "response": "JSON with extracted text and key-value pairs"
        }
    }


# ═══════════════════════════════════════════════════════════════
# ROUTE 3: THE MAIN ENDPOINT — POST /extract
# This is where the real work happens!
# Send an image → Get back structured JSON
# ═══════════════════════════════════════════════════════════════
@app.post(
    "/extract",
    summary="Extract Text from Document Image",
    description="""
    Upload an invoice or document image.
    Returns structured JSON with:
    - All extracted text (full OCR)
    - Key-value pairs (invoice number, date, total, etc.)
    - Document type classification

    Supported formats: JPEG, PNG, TIFF
    Max file size: 20MB
    """
)
async def extract_text_from_image(
    file: UploadFile = File(
        ...,
        description="Upload your invoice or document image (JPEG, PNG, TIFF)"
    )
):
    """
    The main OCR endpoint.

    Flow:
    1. Receive the uploaded image file
    2. Validate the file (correct type, not too large)
    3. Pass to score() for OCI Vision processing
    4. Return structured JSON result
    """

    # ── Track processing time ──────────────────────────────────
    start_time = time.time()
    logger.info(f"📨 Received file: {file.filename} ({file.content_type})")

    # ── Validate file type ─────────────────────────────────────
    # Only allow image files — no videos, PDFs converted to images, etc.
    ALLOWED_TYPES = {
        "image/jpeg", "image/jpg",
        "image/png",
        "image/tiff",
        "image/webp"
    }

    if file.content_type not in ALLOWED_TYPES:
        # If wrong file type, return a nice error message
        raise HTTPException(
            status_code=400,   # 400 = Bad Request
            detail={
                "error":    "Invalid file type",
                "received": file.content_type,
                "allowed":  list(ALLOWED_TYPES),
                "tip":      "Please upload a JPEG, PNG, or TIFF image"
            }
        )

    # ── Read the uploaded file into memory ─────────────────────
    # This reads the image bytes from the HTTP request
    try:
        image_bytes = await file.read()
    except Exception as e:
        raise HTTPException(
            status_code=500,
            detail=f"Failed to read uploaded file: {str(e)}"
        )

    # ── Check file size ────────────────────────────────────────
    MAX_SIZE_MB = 20
    MAX_SIZE_BYTES = MAX_SIZE_MB * 1024 * 1024  # Convert MB to bytes

    if len(image_bytes) > MAX_SIZE_BYTES:
        raise HTTPException(
            status_code=413,   # 413 = Payload Too Large
            detail={
                "error": f"File too large",
                "max_allowed": f"{MAX_SIZE_MB}MB",
                "received":    f"{len(image_bytes) / (1024*1024):.1f}MB"
            }
        )

    logger.info(f"✅ File validation passed. Size: {len(image_bytes)/1024:.1f} KB")

    # ── Call score() with the image bytes ─────────────────────
    # This is where all the magic happens!
    # score() calls OCI Vision and returns structured JSON
    try:
        logger.info("🧠 Calling score() for OCR processing...")
        result = score(image_bytes, input_type="bytes")
    except Exception as e:
        logger.error(f"❌ score() failed: {e}")
        raise HTTPException(
            status_code=500,
            detail={
                "error":   "OCR processing failed",
                "details": str(e)
            }
        )

    # ── Calculate processing time ──────────────────────────────
    processing_time = round(time.time() - start_time, 2)

    # ── Add metadata to the response ──────────────────────────
    result["metadata"] = {
        "filename":        file.filename,
        "file_size_kb":    round(len(image_bytes) / 1024, 1),
        "content_type":    file.content_type,
        "processing_time": f"{processing_time}s",
        "service":         "OCI Vision OCR Inference v1.0"
    }

    logger.info(f"✅ Request completed in {processing_time}s")
    logger.info(f"   Success: {result.get('success')}")

    # Return the result as a JSON response
    return JSONResponse(
        content=result,
        status_code=200 if result.get("success") else 500
    )


# ═══════════════════════════════════════════════════════════════
# ROUTE 4: POST /extract/url — Extract from Image URL (Bonus!)
# ═══════════════════════════════════════════════════════════════

class URLRequest(BaseModel):
    """Request body for URL-based extraction."""
    image_url: str
    description: Optional[str] = None


@app.post(
    "/extract/url",
    summary="Extract Text from Image URL",
    description="Provide an image URL instead of uploading a file"
)
async def extract_from_url(request: URLRequest):
    """
    Alternative endpoint — provide an image URL instead of uploading.
    We download the image and process it the same way.
    """
    import requests as http_requests

    logger.info(f"🌐 Downloading image from URL: {request.image_url}")

    # Download the image from the URL
    try:
        response = http_requests.get(
            request.image_url,
            timeout=30,          # Wait max 30 seconds
            stream=True
        )
        response.raise_for_status()
        image_bytes = response.content
    except Exception as e:
        raise HTTPException(
            status_code=400,
            detail=f"Failed to download image from URL: {str(e)}"
        )

    # Process with score() (same as file upload!)
    result = score(image_bytes, input_type="bytes")
    result["metadata"] = {
        "source":      "url",
        "image_url":   request.image_url,
        "service":     "OCI Vision OCR Inference v1.0"
    }

    return JSONResponse(content=result)


# ═══════════════════════════════════════════════════════════════
# Start the server when we run: python api.py
# ═══════════════════════════════════════════════════════════════
if __name__ == "__main__":
    port = int(os.getenv("APP_PORT", 8000))

    print("\n" + "=" * 60)
    print("🚀 OCI Vision OCR Inference Service Starting...")
    print("=" * 60)
    print(f"   API URL:  http://localhost:{port}")
    print(f"   Docs URL: http://localhost:{port}/docs")
    print(f"   Health:   http://localhost:{port}/health")
    print("=" * 60 + "\n")

    # Start the FastAPI server with Uvicorn
    uvicorn.run(
        "api:app",          # Run the 'app' object in 'api.py'
        host="0.0.0.0",     # Accept connections from any IP
        port=port,
        reload=False,       # Set to True during development for auto-reload
        log_level="info"
    )

📋 Step 4 — Explaining important lines:

  • @app.post("/extract") → This decorator creates an endpoint at URL /extract that listens for POST requests (file uploads)
  • UploadFile = File(...) → Tells FastAPI to accept a file upload in this field. The ... means it's required
  • await file.read() → Read the uploaded file asynchronously (without blocking the server)
  • raise HTTPException(status_code=400) → When something goes wrong, we return a proper HTTP error with a helpful message
  • host="0.0.0.0" → Makes the server accessible from outside the container (important for Docker!)

🔹 SECTION 6 — Dockerfile (Packing the App into a Box 📦)

📋 Step 1 — What this file does:
A Dockerfile is like a recipe for packing our entire app into a portable box (container) 📦. Once packed, this box can run ANYWHERE — on your laptop, on a server, in Oracle Cloud — and it will behave exactly the same way every time!

📋 Step 2 — Why we need it:
Without Docker, sharing our app is hard. You'd need to install Python, all the libraries, and configure everything exactly right. With Docker, you just share the box and it runs — no configuration needed!

🧒 Super simple analogy: Imagine you baked a birthday cake 🎂. Normally you'd have to teach the person all the steps to bake it. With Docker, you just ship them the entire finished cake — perfectly packaged, ready to eat!

# ═══════════════════════════════════════════════════════════════
# Dockerfile — Recipe for packing our OCR service into a box 📦
# ═══════════════════════════════════════════════════════════════

# ─── STAGE 1: Base Image ──────────────────────────────────────
# We start FROM an official Python image — like starting with a
# clean kitchen that already has the oven and basic tools!
# "slim-bookworm" = small size + Debian Linux base
FROM python:3.11-slim-bookworm

# ─── Set the author label ─────────────────────────────────────
LABEL maintainer="your-name@email.com"
LABEL description="OCI Vision OCR Inference Service"
LABEL version="1.0.0"

# ─── STAGE 2: System-level setup ──────────────────────────────
# Install system tools that Python packages need
# These are like installing kitchen equipment before cooking!
RUN apt-get update && apt-get install -y \
    # These are needed for Pillow (image processing library):
    libpq-dev           \   
    gcc                 \
    # Image format support:
    libjpeg-dev         \
    libpng-dev          \
    libtiff-dev         \
    libwebp-dev         \
    # Clean up apt cache to keep the image small:
    && apt-get clean    \
    && rm -rf /var/lib/apt/lists/*

# ─── STAGE 3: Working Directory ───────────────────────────────
# Create and set the working directory inside the container
# This is like creating a dedicated workspace folder!
WORKDIR /app

# ─── STAGE 4: Install Python dependencies ─────────────────────
# Copy requirements.txt FIRST (before copying code)
# Why? Docker is smart — if requirements.txt hasn't changed,
# it skips this step. This makes rebuilds much faster! ⚡
COPY requirements.txt .

# Install all our Python libraries from the shopping list
RUN pip install --no-cache-dir --upgrade pip \
    && pip install --no-cache-dir -r requirements.txt

# ─── STAGE 5: Copy our application code ───────────────────────
# Now copy all our files into the container
# We do this AFTER installing requirements (for Docker caching)
COPY score.py    .    
COPY api.py      .
COPY .env        .

# ─── STAGE 6: Copy OCI credentials ───────────────────────────
# Copy the OCI API private key file into the container
# This is needed to authenticate with Oracle Cloud
COPY oci_api_key.pem .

# ─── STAGE 7: Security — Don't run as root! ───────────────────
# Create a non-root user for security
# It's safer to run apps as a regular user, not the administrator!
RUN useradd --create-home --shell /bin/bash appuser \
    && chown -R appuser:appuser /app
USER appuser

# ─── STAGE 8: Expose the port ─────────────────────────────────
# Tell Docker that our app uses port 8000
# Like labeling the front door of our house!
EXPOSE 8000

# ─── STAGE 9: Health Check ────────────────────────────────────
# Docker will check if our service is healthy every 30 seconds
# If it fails 3 times, Docker knows something is wrong
HEALTHCHECK --interval=30s --timeout=10s --start-period=10s --retries=3 \
    CMD python -c "import requests; requests.get('http://localhost:8000/health')" \
    || exit 1

# ─── STAGE 10: Start Command ──────────────────────────────────
# This is the command that runs when the container starts
# Like pressing the "Power On" button!
CMD ["python", "api.py"]

📋 Step 4 — Explaining important lines:

  • FROM python:3.11-slim-bookworm → Start with a small official Python container (our blank canvas)
  • WORKDIR /app → All our files live inside /app folder in the container
  • COPY requirements.txt . → Copy BEFORE code (for Docker layer caching — speeds up rebuilds!)
  • RUN pip install --no-cache-dir → Install libraries WITHOUT saving download cache (keeps container small)
  • USER appuser → Security best practice! Never run apps as root inside containers
  • EXPOSE 8000 → Tells Docker which port to "open a window" on
  • HEALTHCHECK → Docker regularly pings our health endpoint to make sure we're alive
  • CMD ["python", "api.py"] → The starting command when container launches

How to Build and Run the Docker Container

# ── Step 1: Build the Docker image ────────────────────────────
# The -t flag gives our image a name and tag
# The . means "use the Dockerfile in the current folder"
docker build -t oci-vision-ocr:v1.0 .

# You'll see Docker executing each step from our Dockerfile!
# This takes 1-3 minutes the first time (downloading libraries)

# ── Step 2: Run the container ─────────────────────────────────
# -d = run in background (detached mode)
# -p 8000:8000 = map port 8000 on your computer to port 8000 in container
# --name = give the container a friendly name
docker run -d \
  -p 8000:8000 \
  --name ocr-service \
  oci-vision-ocr:v1.0

# ── Step 3: Check it's running ────────────────────────────────
docker ps
# You should see your container in the list with STATUS: Up

# ── Step 4: Check the logs ────────────────────────────────────
docker logs ocr-service
# Should show: "🚀 OCI Vision OCR Inference Service Starting..."

# ── Step 5: Test the health endpoint ─────────────────────────
curl http://localhost:8000/health
# Should return: {"status": "healthy", ...}

# ── Step 6: Stop the container ────────────────────────────────
docker stop ocr-service
docker rm ocr-service

🔹 SECTION 7 — Execution Flow (The Complete Journey 🗺️)

Let's trace exactly what happens when someone sends us an invoice image. Think of it like following a parcel from your door to the delivery company and back!

┌──────────────────────────────────────────────────────────────────┐
│              Complete Request Flow — Step by Step                │
└──────────────────────────────────────────────────────────────────┘

  👤 USER sends a request
        │
        │  HTTP POST /extract
        │  Body: invoice.jpg (image file)
        │
        ▼
  ┌─────────────────────────────────┐
  │  api.py — FastAPI Endpoint      │
  │                                 │
  │  1. Receive the file            │
  │  2. Validate type + size        │
  │  3. Read bytes from request     │
  └──────────────┬──────────────────┘
                 │
                 │  Calls score(image_bytes, "bytes")
                 │
                 ▼
  ┌─────────────────────────────────┐
  │  score.py — The Brain           │
  │                                 │
  │  1. Load OCIConfig (settings)   │
  │  2. Create OCREngine            │
  │  3. Convert image → Base64      │
  │  4. Call OCI Vision API  ────── │──────────────────┐
  │  5. Get response back    ◄───── │──────────────────┤
  │  6. Extract full text           │                  │
  │  7. Extract key-value pairs     │                  ▼
  │  8. Run pattern matching        │   ☁️  OCI Vision AI
  │  9. Build final JSON            │   (Oracle Cloud)
  └──────────────┬──────────────────┘   Reads the image,
                 │                      finds text blocks,
                 │  Returns JSON result  detects key-value
                 │                      pairs, classifies
                 ▼                      document type
  ┌─────────────────────────────────┐
  │  api.py — Returns Response      │
  │                                 │
  │  Adds metadata (filename, time) │
  │  Returns HTTP 200 + JSON body   │
  └──────────────┬──────────────────┘
                 │
                 ▼
        👤 USER receives structured JSON!

  Total time: typically 2-5 seconds ⚡

Data Transformation at Each Step

  📷 Input:   invoice.jpg (binary file — raw image data)
       │
       ▼
  🔤 score.py: Convert to Base64 string
               "iVBORw0KGgoAAAANSUhEUgAA..." (long text)
       │
       ▼
  ☁️ OCI Vision: Returns OCR result (raw API response)
               pages[0].lines[0].text = "INVOICE #1234"
               pages[0].lines[1].text = "Date: 2026-04-18"
               key_value_fields[0].name = "Invoice Number"
               key_value_fields[0].value = "INV-1234"
       │
       ▼
  🧹 score.py: Process and organize into clean JSON
       │
       ▼
  📋 Output:  {
                "success": true,
                "document": { "type": "INVOICE" },
                "extracted_text": { "full_text": "INVOICE #1234 ..." },
                "key_values": { ... },
                "summary": { "invoice_number": "1234", ... }
              }

🔹 SECTION 8 — Sample Input and Output

Now let's see what our service actually looks like when it's working! This is the exciting part — real examples! 

Sample Request — How to Call Our API

📋 Using curl (command line):

# Send an invoice image to our OCR service
# -X POST = use POST method
# -F "file=@invoice.jpg" = upload file named "invoice.jpg"
curl -X POST \
  http://localhost:8000/extract \
  -F "file=@invoice.jpg" \
  -H "accept: application/json"

📋 Using Python (requests library):

# test_api.py — A simple test script to call our OCR API

import requests
import json

# ── Define the API endpoint ───────────────────────────────────
API_URL = "http://localhost:8000/extract"

# ── Open and upload the invoice image ─────────────────────────
image_path = "sample_invoice.jpg"

print(f"📷 Uploading image: {image_path}")
print(f"🌐 Sending to: {API_URL}")
print("-" * 40)

# Open the image file and send it to our API
with open(image_path, "rb") as image_file:
    response = requests.post(
        API_URL,
        files={"file": (image_path, image_file, "image/jpeg")}
    )

# ── Check if request was successful ───────────────────────────
print(f"✅ Status Code: {response.status_code}")

# ── Print the result nicely ───────────────────────────────────
if response.status_code == 200:
    result = response.json()
    print("\n📋 EXTRACTED DATA:")
    print(json.dumps(result, indent=2))
else:
    print(f"❌ Error: {response.text}")

Sample Input — What an Invoice Looks Like

Imagine our input image shows this invoice:

┌─────────────────────────────────────────┐
│           ACME CORP                     │
│        123 Business Street              │
│        Mumbai, MH 400001                │
│        Email: billing@acme.com          │
│                                         │
│  INVOICE                                │
│  Invoice No:   INV-2026-0042            │
│  Date:         18-Apr-2026              │
│  Due Date:     18-May-2026              │
│                                         │
│  Bill To:                               │
│  Tech Startup Ltd                       │
│  456 Startup Road, Bangalore            │
│                                         │
│  ─────────────────────────────────────  │
│  Item          Qty    Price    Total    │
│  ─────────────────────────────────────  │
│  Web Design    1      $800     $800     │
│  Hosting       12     $50      $600     │
│  Support       5 hr   $100     $500     │
│  ─────────────────────────────────────  │
│  Subtotal:                    $1,900    │
│  Tax (18% GST):                 $342    │
│  Grand Total:                 $2,242    │
└─────────────────────────────────────────┘

Sample Output — What Our Service Returns

{
  "success": true,

  "document": {
    "type": "INVOICE",
    "type_confidence": "97.3%",
    "total_characters": 482,
    "total_lines": 24
  },

  "extracted_text": {
    "full_text": "ACME CORP\n123 Business Street\nMumbai, MH 400001\nEmail: billing@acme.com\n
                 INVOICE\nInvoice No: INV-2026-0042\nDate: 18-Apr-2026\nDue Date: 18-May-2026\n
                  Bill To:\nTech Startup Ltd\n456 Startup Road, Bangalore\nItem Qty Price Total\n
                  Web Design 1 $800 $800\nHosting 12 $50 $600\n
                 Support 5 hr $100 $500\nSubtotal: $1,900\nTax (18% GST): $342\nGrand Total: $2,242", "line_count": 18 }, "key_values": { "oci_detected": { "invoice_number": { "value": "INV-2026-0042", "confidence": 98.1 }, "invoice_date": { "value": "18-Apr-2026", "confidence": 96.5 }, "due_date": { "value": "18-May-2026", "confidence": 95.8 }, "total_amount": { "value": "$2,242", "confidence": 97.2 }, "vendor_name": { "value": "ACME CORP", "confidence": 94.3 }, "subtotal": { "value": "$1,900", "confidence": 93.1 }, "tax": { "value": "$342", "confidence": 91.7 } }, "pattern_extracted": { "invoice_number": "INV-2026-0042", "date": "18-Apr-2026", "total_amount": "2,242", "tax_amount": "342", "vendor_name": "ACME CORP", "email": "billing@acme.com" } }, "summary": { "invoice_number": "INV-2026-0042", "date": "18-Apr-2026", "due_date": "18-May-2026", "total_amount": "$2,242", "tax_amount": "$342", "vendor_name": "ACME CORP", "email": "billing@acme.com" }, "metadata": { "filename": "sample_invoice.jpg", "file_size_kb": 145.2, "content_type": "image/jpeg", "processing_time": "2.84s", "service": "OCI Vision OCR Inference v1.0" } }

Look at that! We sent an image and got back:

  • ✅ The document type (INVOICE, detected with 97.3% confidence!)
  • ✅ All the raw text that was read from the image
  • ✅ Structured key-value pairs (invoice number, date, total, etc.)
  • ✅ A clean summary with the most important fields
  • ✅ Metadata (how long it took, file size, etc.)

📁 Final Project Structure

Here is how your complete project folder should look:

oci-vision-ocr-service/
│
├── 📄 score.py              ← The brain (main inference logic)
├── 📄 api.py                ← FastAPI web server
├── 📄 requirements.txt      ← Shopping list of dependencies
├── 📄 Dockerfile            ← Container packaging recipe
│
├── 📄 .env                  ← Secret settings (NEVER commit to Git!)
├── 📄 .gitignore            ← Tells Git to ignore .env and key files
│
├── 🔑 oci_api_key.pem       ← OCI private key (NEVER commit to Git!)
│
├── 📄 test_api.py           ← Quick test script
│
└── 📄 sample_invoice.jpg    ← Sample image for testing

📄 Create your .gitignore file:

# .gitignore — Tell Git to ignore sensitive files!

# Secret settings — NEVER push these to GitHub!
.env
*.pem
*.key

# Python temporary files
__pycache__/
*.pyc
*.pyo
*.pyd
.Python
*.egg-info/

# Virtual environment
venv/
env/
.venv/

# Test output files
test_output/
*.log

How to Run Everything — Step by Step

# ── OPTION A: Run directly with Python ────────────────────────

# Step 1: Create a virtual environment (isolated Python space)
python -m venv venv

# Step 2: Activate it
# On Mac/Linux:
source venv/bin/activate
# On Windows:
venv\Scripts\activate

# Step 3: Install dependencies
pip install -r requirements.txt

# Step 4: Test score.py directly
python score.py sample_invoice.jpg

# Step 5: Start the API server
python api.py
# Server starts at: http://localhost:8000

# Step 6: Open your browser and visit:
# http://localhost:8000/docs  ← Interactive API docs!
# http://localhost:8000/health ← Health check

# ── OPTION B: Run with Docker ──────────────────────────────────

# Step 1: Build the Docker image
docker build -t oci-vision-ocr:v1.0 .

# Step 2: Run the container
docker run -d -p 8000:8000 --name ocr-service oci-vision-ocr:v1.0

# Step 3: Test it
curl -X POST http://localhost:8000/extract -F "file=@sample_invoice.jpg"

🏆 Quick Summary — What We Built

We built a complete ML Inference Service from scratch! Here's everything:

  • score.py → The brain! Connects to OCI Vision, sends the image, extracts text and key-value pairs, and returns structured JSON. It uses both OCI's built-in detection AND our own pattern matching for maximum accuracy!
  • api.py → The front door! A FastAPI web server with endpoints for file upload, URL-based processing, and health checks. Has Swagger docs automatically at /docs!
  • requirements.txt → The shopping list of all Python libraries we need. One command to install everything!
  • Dockerfile → The packaging recipe. Puts our entire app into a portable Docker container that runs anywhere!
  • Security → Private keys and secrets stored in .env file, non-root Docker user, file validation, size limits — all the safety checks!

Happy coding! 🤖✨

Comments