A real-world AI system that can read invoices, detect fraud, and explain its decisions in plain English — all automatically!
Think of it like building a smart bank employee 🏦 who:
- 📷 Reads any invoice photo automatically
- 🔍 Checks if the claim amount is honest
- 🤖 Uses an AI brain (LLM) to explain WHY something looks suspicious
- 🌐 Exposes everything as an API so any app can use it
💡 Real-world use: Insurance companies, finance teams, and audit departments use systems exactly like this to catch fraud every single day!
🗺️ Complete System Architecture
┌─────────────────────────────────────────────────────────────────────────┐
│ Invoice Compliance & Fraud Detection — Full Flow │
└─────────────────────────────────────────────────────────────────────────┘
👤 User uploads invoice + submits claim amount
│
▼
┌───────────────────┐
│ FastAPI (app.py)│ ← Web API — the front door
│ POST /upload │
│ POST /validate │
└────────┬──────────┘
│
┌──────┴──────────────────────────────────────┐
│ │
▼ ▼
┌─────────────────────┐ ┌──────────────────────────┐
│ OCI Object Storage │ │ OCI Document Understanding│
│ (stores invoice) │────────►│ (reads invoice, extracts │
└─────────────────────┘ │ amount, vendor, date) │
└────────────┬─────────────┘
│
▼
┌────────────────────────┐
│ Rule Engine │
│ claim > invoice? 🚩 │
│ claim > policy? 🚩 │
└────────────┬───────────┘
│
▼
┌────────────────────────┐
│ Hugging Face LLM │
│ (Flan-T5) │
│ Explains the verdict │
└────────────┬───────────┘
│
▼
{
"status": "Suspicious",
"reason": "Claim of $800
exceeds invoice
total of $500"
}
🔹 PHASE 1 — Understanding Invoice Fraud (Simple Examples!)
Before we build anything, let's understand WHAT we are protecting against. Think of it like understanding the rules of a game before you play! 🎮
What is Invoice Fraud?
Imagine your school gives you $50 to buy art supplies. You buy them for $30 but claim you spent $50 and pocket the $20 difference. That's fraud! 😤
In real business, people sometimes:
- 🚩 Inflate claims — Submit a claim for $1,000 when the invoice shows $600
- 🚩 Fake invoices — Create invoices for work that was never done
- 🚩 Exceed policy limits — Claim $5,000 when the policy only covers $3,000
- 🚩 Wrong dates — Submit invoices for services outside the policy period
Our Two Fraud Rules (Keep It Simple!)
Rule 1: If claim_amount > invoice_total → 🚩 SUSPICIOUS
"You can't claim more than what you actually paid!"
Rule 2: If claim_amount > policy_limit → 🚩 SUSPICIOUS
"Your insurance only covers up to a certain amount!"
If neither rule triggers → ✅ VALID
Our LLM (AI brain) then writes a clear English explanation of WHY we made that decision!
🔹 PHASE 2 — Architecture Deep Dive
Let's tell the story of what happens when a user uploads an invoice. Read it like a story! 📖
STORY: "The Journey of an Invoice"
📸 Chapter 1 — Upload
Maria uploads her invoice photo to our API.
Our system stores it safely in OCI Object Storage.
(Like putting a letter in a very safe filing cabinet! 🗄️)
🤖 Chapter 2 — Reading
OCI Document Understanding (our AI reader) looks at the photo.
It reads: Vendor = "ACME Corp", Total = $500, Date = 2026-04-01
(Like a super-fast, super-accurate typist who reads any document!)
📋 Chapter 3 — Rules Check
Maria claims $800 from insurance.
Rule Engine checks: $800 > $500 invoice? YES → Flag it! 🚩
(Like a teacher checking if your homework answer matches the key!)
🧠 Chapter 4 — LLM Explanation
We ask Flan-T5: "Is this claim valid? Explain why."
LLM writes: "The claim of $800 exceeds the invoice total of $500..."
(Like asking a wise professor to explain the teacher's decision!)
📤 Chapter 5 — Response
Our API returns: { "status": "Suspicious", "reason": "..." }
Maria sees exactly why her claim was flagged! ✅
Project Files Overview
fraud-detection-system/ │ ├── 📄 app.py ← Brain + API (FastAPI + all logic) ├── 📄 ocr_engine.py ← OCI Document Understanding logic ├── 📄 rule_engine.py ← Fraud detection rules ├── 📄 llm_engine.py ← Hugging Face Flan-T5 logic ├── 📄 requirements.txt ← Shopping list of tools ├── 📄 Dockerfile ← Packaging box ├── 📄 .env ← Secret settings └── 📄 .gitignore ← Protect secrets from Git
🔹 PHASE 3 — OCI Setup
Step A: Get OCI Credentials
Our system needs to talk to Oracle Cloud. For that, we need a special key — like a door key 🔑 that unlocks Oracle's services.
- Log in to
cloud.oracle.com - Click your Profile icon → My Profile
- Click API Keys → Add API Key → Generate API Key Pair
- Download the private key file (save as
oci_api_key.pem) - Copy the config snippet and save it
Step B: Create OCI Object Storage Bucket
Our invoices need a safe home in the cloud before we can process them. Let's create a bucket (like a folder in the cloud ☁️):
- OCI Console → Storage → Object Storage → Create Bucket
- Name it:
invoice-fraud-bucket - Visibility: Private
- Click Create
Step C: Create the .env File
📋 Step 1 — What this does: Stores all our secret settings in one safe place. Like a private diary that only our app can read!
📋 Step 2 — Why we need it: We should NEVER hardcode passwords or keys inside our Python code. If we push code to GitHub, our secrets would be exposed to the world!
# .env — Secret settings for our Fraud Detection System # ⚠️ NEVER share this file or push it to GitHub! # ─── OCI Authentication ──────────────────────────────────────── OCI_USER_OCID=ocid1.user.oc1..aaaaaaaXXXXXXXXXXXXX OCI_FINGERPRINT=aa:bb:cc:dd:ee:ff:11:22:33:44 OCI_TENANCY_OCID=ocid1.tenancy.oc1..aaaaaaaXXXXXXXXXXXXX OCI_REGION=us-ashburn-1 OCI_KEY_FILE=./oci_api_key.pem # ─── OCI Resources ───────────────────────────────────────────── COMPARTMENT_ID=ocid1.compartment.oc1..aaaaaaaXXXXXXXXXXXXX BUCKET_NAME=invoice-fraud-bucket OCI_NAMESPACE=your-tenancy-namespace-here # ─── Policy Settings ─────────────────────────────────────────── # Maximum amount insurance will pay (policy limit) POLICY_LIMIT=5000.00 # ─── Hugging Face Settings ───────────────────────────────────── # Which LLM model to use (flan-t5-base is a good balance) HF_MODEL_NAME=google/flan-t5-base # ─── Server Settings ─────────────────────────────────────────── APP_PORT=8000
📋 Step 4 — Key lines explained:
POLICY_LIMIT=5000.00→ The maximum the insurance will pay. Any claim above this gets flagged.HF_MODEL_NAME=google/flan-t5-base→ Which Hugging Face AI model we use to generate explanationsBUCKET_NAME=invoice-fraud-bucket→ The OCI storage bucket we created in Step B
🔹 PHASE 4 — Upload Invoice to OCI Object Storage
📋 Step 1 — What this code will do:
This module handles uploading an invoice file to OCI Object Storage. Think of it like a post office worker 📮 who takes your letter (invoice) and puts it safely in the cloud filing cabinet so other parts of our system can access it!
📋 Step 2 — Why we need it:
OCI Document Understanding service reads documents from Object Storage. So first, we upload the invoice, then we tell OCI "hey, please read the file at this location".
# ═══════════════════════════════════════════════════════════════
# storage_engine.py — Handles uploading files to OCI Object Storage
# ═══════════════════════════════════════════════════════════════
import os
import uuid
import logging
import oci
import oci.object_storage
from dotenv import load_dotenv
load_dotenv()
logger = logging.getLogger(__name__)
class StorageEngine:
"""
Handles uploading invoice files to OCI Object Storage.
Think of this as the filing clerk 📁 who stores all invoices safely!
"""
def __init__(self):
# ── Build OCI config from environment variables ────────
self.oci_config = {
"user": os.getenv("OCI_USER_OCID"),
"fingerprint": os.getenv("OCI_FINGERPRINT"),
"tenancy": os.getenv("OCI_TENANCY_OCID"),
"region": os.getenv("OCI_REGION"),
"key_file": os.getenv("OCI_KEY_FILE"),
}
# ── OCI Object Storage settings ────────────────────────
self.namespace = os.getenv("OCI_NAMESPACE")
self.bucket_name = os.getenv("BUCKET_NAME", "invoice-fraud-bucket")
# ── Create the OCI Object Storage client ───────────────
# This is our "connection" to Oracle Cloud Storage
self.storage_client = oci.object_storage.ObjectStorageClient(
config=self.oci_config
)
logger.info("✅ StorageEngine ready — connected to OCI Object Storage")
def upload_invoice(self, file_bytes: bytes, original_filename: str) -> dict:
"""
Upload an invoice file to OCI Object Storage.
Returns a dict with the object name (so we can reference it later).
Like getting a ticket receipt 🎫 when you check luggage at the airport!
"""
# ── Generate a unique filename to avoid conflicts ──────
# uuid4() generates a random ID like: a1b2c3d4-e5f6-...
# We combine it with the original name for uniqueness
unique_id = str(uuid.uuid4())[:8]
object_name = f"invoices/{unique_id}_{original_filename}"
logger.info(f"📤 Uploading invoice: {object_name} to bucket: {self.bucket_name}")
try:
# ── Upload the file to Object Storage ─────────────
# put_object = "put this file into the bucket"
response = self.storage_client.put_object(
namespace_name=self.namespace,
bucket_name=self.bucket_name,
object_name=object_name,
put_object_body=file_bytes, # The actual file bytes
content_type="application/octet-stream"
)
logger.info(f"✅ Invoice uploaded successfully! Object: {object_name}")
return {
"success": True,
"object_name": object_name,
"bucket": self.bucket_name,
"namespace": self.namespace,
"message": f"Invoice stored as: {object_name}"
}
except Exception as e:
logger.error(f"❌ Upload failed: {e}")
return {
"success": False,
"error": str(e)
}
📋 Step 4 — Key lines explained:
uuid.uuid4()[:8]→ Generates a short random ID so two invoices with the same filename never clashobject_name = f"invoices/{unique_id}_{original_filename}"→ Stores files in an "invoices/" folder inside the bucketput_object(...)→ OCI SDK's method to upload a file — like a "save" button for the cloud!
🔹 PHASE 5 — OCR: Extract Invoice Data
📋 Step 1 — What this code will do:
This is the "reader robot" 🤖. It connects to OCI Document Understanding, sends the uploaded invoice, and extracts three key things: vendor name, invoice total, and invoice date. Like having a very smart assistant who reads every invoice and fills in a form for you!
📋 Step 2 — Why we need it:
We cannot manually read invoices — there could be thousands of them! OCI Document Understanding has pre-trained AI models specifically for invoices that can extract 28+ fields automatically.
# ═══════════════════════════════════════════════════════════════
# ocr_engine.py — Reads invoice data using OCI Document Understanding
# ═══════════════════════════════════════════════════════════════
import os
import re
import base64
import logging
import oci
import oci.ai_document
from dotenv import load_dotenv
from typing import Optional
load_dotenv()
logger = logging.getLogger(__name__)
class OCREngine:
"""
Uses OCI Document Understanding to extract key fields from invoices.
🧒 Think of this as a robot with super reading glasses 👓 that can
look at any invoice and instantly find the important numbers and names!
"""
def __init__(self):
# ── OCI Config ─────────────────────────────────────────
self.oci_config = {
"user": os.getenv("OCI_USER_OCID"),
"fingerprint": os.getenv("OCI_FINGERPRINT"),
"tenancy": os.getenv("OCI_TENANCY_OCID"),
"region": os.getenv("OCI_REGION"),
"key_file": os.getenv("OCI_KEY_FILE"),
}
self.compartment_id = os.getenv("COMPARTMENT_ID")
self.namespace = os.getenv("OCI_NAMESPACE")
self.bucket_name = os.getenv("BUCKET_NAME")
# ── Create OCI Document Understanding client ───────────
self.doc_client = oci.ai_document.AIServiceDocumentClient(
config=self.oci_config
)
logger.info("✅ OCREngine ready — OCI Document Understanding connected")
def extract_invoice_data(
self,
object_name: str,
file_bytes: Optional[bytes] = None
) -> dict:
"""
Extract key fields from an invoice stored in OCI Object Storage.
We call OCI's pre-trained INVOICE key-value extraction model.
It automatically finds: invoice_id, vendor_name, total, date, etc.
Returns a clean dictionary with the extracted fields!
"""
logger.info(f"🔍 Extracting data from invoice: {object_name}")
try:
# ── METHOD: Use Object Storage location ───────────
# We tell OCI: "The invoice is in this bucket at this path"
input_location = oci.ai_document.models.ObjectStorageLocations(
object_locations=[
oci.ai_document.models.ObjectLocation(
namespace_name=self.namespace,
bucket_name=self.bucket_name,
object_name=object_name
)
]
)
# ── Output location (where OCI stores raw results) ─
# OCI Document Understanding writes raw JSON results to a bucket
output_location = oci.ai_document.models.OutputLocation(
namespace_name=self.namespace,
bucket_name=self.bucket_name,
prefix="ocr-results"
)
# ── Define what we want OCI to extract ─────────────
# KEY_VALUE_DETECTION = find field:value pairs like an invoice form
key_value_feature = oci.ai_document.models.DocumentKeyValueExtractionFeature()
# ── Create the Document Processing Job ────────────
# Think of this like submitting a job to a printing queue!
processor_config = oci.ai_document.models.GeneralProcessorConfig(
features=[key_value_feature],
document_type="INVOICE" # Tell OCI this is an invoice!
)
# ── Create the actual processor job ────────────────
composite_client = oci.ai_document.AIServiceDocumentClientCompositeOperations(
self.doc_client
)
import uuid
job_details = oci.ai_document.models.CreateProcessorJobDetails(
display_name=f"fraud-detect-{str(uuid.uuid4())[:8]}",
compartment_id=self.compartment_id,
input_location=input_location,
output_location=output_location,
processor_config=processor_config
)
# ── Wait for the job to complete ───────────────────
# create_processor_job_and_wait_for_state = run AND wait!
logger.info("⏳ Running OCI Document Understanding job...")
response = composite_client.create_processor_job_and_wait_for_state(
create_processor_job_details=job_details,
wait_for_states=[
oci.ai_document.models.ProcessorJob.LIFECYCLE_STATE_SUCCEEDED,
oci.ai_document.models.ProcessorJob.LIFECYCLE_STATE_FAILED
]
)
if response.data.lifecycle_state == "FAILED":
raise Exception("OCI Document Understanding job failed!")
logger.info("✅ OCI Document Understanding job succeeded!")
# ── Read the results from Object Storage ───────────
# OCI writes the results as JSON to the output_location bucket
import oci.object_storage
storage_client = oci.object_storage.ObjectStorageClient(
config=self.oci_config
)
# The result file path format from OCI
result_prefix = f"ocr-results/{response.data.id}"
result_object = f"{result_prefix}/results.json"
raw_result = storage_client.get_object(
namespace_name=self.namespace,
bucket_name=self.bucket_name,
object_name=result_object
)
import json
result_json = json.loads(raw_result.data.content.decode("utf-8"))
# ── Parse the results into clean format ────────────
return self._parse_invoice_fields(result_json)
except Exception as e:
logger.error(f"❌ OCR extraction failed: {e}")
# Return empty fields so system can still run with defaults
return self._empty_invoice_fields(str(e))
def _parse_invoice_fields(self, result_json: dict) -> dict:
"""
Parse OCI Document Understanding raw JSON into clean fields.
OCI returns a complex nested JSON. We pick out just what we need:
- invoice_total → The total amount on the invoice
- vendor_name → Who issued the invoice
- invoice_date → When the invoice was issued
- invoice_id → The invoice number
"""
extracted = {
"invoice_total": None,
"vendor_name": None,
"invoice_date": None,
"invoice_id": None,
"raw_fields": {},
"success": True,
"error": None
}
try:
# Navigate OCI's JSON structure to find pages and fields
pages = result_json.get("pages", [])
for page in pages:
# Each page has document_fields with key-value pairs
doc_fields = page.get("documentFields", [])
for field in doc_fields:
field_name = field.get("fieldLabel", {}).get("name", "").lower()
field_value = field.get("fieldValue", {}).get("text", "")
confidence = field.get("confidence", 0)
# Store all raw fields for reference
extracted["raw_fields"][field_name] = {
"value": field_value,
"confidence": round(confidence * 100, 1)
}
# Map OCI field names → our clean field names
if "total" in field_name and confidence > 0.5:
extracted["invoice_total"] = self._clean_amount(field_value)
elif "vendor" in field_name and confidence > 0.5:
extracted["vendor_name"] = field_value.strip()
elif "date" in field_name and "invoice" in field_name and confidence > 0.5:
extracted["invoice_date"] = field_value.strip()
elif "invoice" in field_name and "id" in field_name:
extracted["invoice_id"] = field_value.strip()
except Exception as e:
logger.warning(f"⚠️ Parsing warning: {e}")
extracted["error"] = str(e)
logger.info(f"📊 Parsed fields: total={extracted['invoice_total']}, "
f"vendor={extracted['vendor_name']}, date={extracted['invoice_date']}")
return extracted
def _clean_amount(self, amount_str: str) -> Optional[float]:
"""
Convert a messy amount string like "$1,500.00" to a clean float 1500.0.
Like a math teacher cleaning up a messy calculation! 📐
"""
if not amount_str:
return None
try:
# Remove currency symbols, commas, spaces → keep numbers and dot
cleaned = re.sub(r'[^\d.]', '', amount_str.replace(',', ''))
return float(cleaned) if cleaned else None
except (ValueError, AttributeError):
return None
def _empty_invoice_fields(self, error_msg: str) -> dict:
"""Return empty fields when extraction fails — prevents crashes."""
return {
"invoice_total": None,
"vendor_name": "Unknown",
"invoice_date": "Unknown",
"invoice_id": "Unknown",
"raw_fields": {},
"success": False,
"error": error_msg
}
📋 Step 4 — Key lines explained:
document_type="INVOICE"→ Tells OCI to use its pre-trained INVOICE model — knows all 28+ invoice fields!create_processor_job_and_wait_for_state(...)→ Submits the reading job AND waits for it to finish automaticallyLIFECYCLE_STATE_SUCCEEDED→ We wait until the job says "I'm done!" before reading results_clean_amount("$1,500.00")→ Removes $ signs and commas so we can do math on it
🔹 PHASE 6 — Rule Engine (Fraud Detection Rules! 🚨)
📋 Step 1 — What this code will do:
This is our fraud detective! 🕵️ It checks two simple rules and produces a verdict. Like a teacher checking your exam answers against the answer key!
📋 Step 2 — Why we need it:
Rules are deterministic (always give the same answer for the same input). They are fast, explainable, and reliable. The LLM gives us natural language — the rule engine gives us the hard facts!
# ═══════════════════════════════════════════════════════════════
# rule_engine.py — The Fraud Detection Rule Engine
# ═══════════════════════════════════════════════════════════════
# Think of this like a strict teacher who checks your work
# against a very clear answer key! 📝
import os
import logging
from dataclasses import dataclass, field
from typing import List, Optional
from dotenv import load_dotenv
load_dotenv()
logger = logging.getLogger(__name__)
@dataclass
class RuleViolation:
"""
Represents a single rule that was broken.
Like a note the teacher writes when you get an answer wrong!
"""
rule_name: str # Name of the rule that was broken
description: str # Simple explanation of why
severity: str # "HIGH" or "MEDIUM" — how serious is it?
flag_value: float # The number that triggered the flag
limit_value: float # The limit that was exceeded
@dataclass
class RuleEngineResult:
"""
Complete result from the rule engine check.
Like a report card 📊 showing all the checks we did!
"""
# Overall verdict
is_suspicious: bool # True if ANY rule was broken
status: str # "Valid" or "Suspicious"
violations: List[RuleViolation] = field(default_factory=list)
# The numbers we checked
claim_amount: Optional[float] = None
invoice_total: Optional[float] = None
policy_limit: Optional[float] = None
# Summary for LLM to use
rule_summary: str = ""
class RuleEngine:
"""
Applies fraud detection rules to an invoice + claim.
Rules implemented:
─────────────────────────────────────────
Rule 1: claim_amount > invoice_total → 🚩 Flag!
"You can't claim more than you actually paid!"
Rule 2: claim_amount > policy_limit → 🚩 Flag!
"Policy doesn't cover amounts this large!"
Rule 3: invoice_total is missing/zero → ⚠️ Warn!
"We couldn't read the invoice amount — suspicious!"
─────────────────────────────────────────
"""
def __init__(self):
# Load policy limit from environment settings
self.policy_limit = float(os.getenv("POLICY_LIMIT", "5000.0"))
logger.info(f"✅ RuleEngine ready — Policy limit: ${self.policy_limit:,.2f}")
def check(
self,
claim_amount: float,
invoice_data: dict,
custom_policy: Optional[float] = None
) -> RuleEngineResult:
"""
Run all fraud detection rules against the claim + invoice.
Arguments:
claim_amount → Amount the user wants to claim
invoice_data → Dictionary from OCREngine (has invoice_total, etc.)
custom_policy → Optional override for policy limit
Returns:
RuleEngineResult with verdict + all details
"""
# Use custom policy limit if provided, otherwise use default
policy_limit = custom_policy if custom_policy else self.policy_limit
invoice_total = invoice_data.get("invoice_total")
violations = []
logger.info(f"⚖️ Checking: claim=${claim_amount}, "
f"invoice=${invoice_total}, policy=${policy_limit}")
# ══════════════════════════════════════════════════════
# RULE 1: Claim must not exceed invoice total
# Like: "You can't claim more than your receipt shows!"
# ══════════════════════════════════════════════════════
if invoice_total is not None:
if claim_amount > invoice_total:
violation = RuleViolation(
rule_name="CLAIM_EXCEEDS_INVOICE",
description=(
f"The claim amount ${claim_amount:,.2f} exceeds the "
f"invoice total ${invoice_total:,.2f} by "
f"${(claim_amount - invoice_total):,.2f}. "
f"A claim cannot be larger than what was actually paid."
),
severity="HIGH",
flag_value=claim_amount,
limit_value=invoice_total
)
violations.append(violation)
logger.warning(f"🚩 Rule 1 VIOLATED: claim ${claim_amount} > invoice ${invoice_total}")
else:
logger.info(f"✅ Rule 1 PASSED: claim ${claim_amount} ≤ invoice ${invoice_total}")
else:
# Invoice total could not be read — suspicious by itself!
violation = RuleViolation(
rule_name="INVOICE_AMOUNT_UNREADABLE",
description=(
"The invoice total amount could not be extracted from the document. "
"This may indicate a fraudulent, altered, or illegible invoice."
),
severity="MEDIUM",
flag_value=claim_amount,
limit_value=0
)
violations.append(violation)
logger.warning("⚠️ Invoice total not readable — adding as suspicious indicator")
# ══════════════════════════════════════════════════════
# RULE 2: Claim must not exceed policy limit
# Like: "Insurance only covers up to a certain amount!"
# ══════════════════════════════════════════════════════
if claim_amount > policy_limit:
violation = RuleViolation(
rule_name="CLAIM_EXCEEDS_POLICY",
description=(
f"The claim amount ${claim_amount:,.2f} exceeds the "
f"policy limit of ${policy_limit:,.2f} by "
f"${(claim_amount - policy_limit):,.2f}. "
f"This claim is not eligible for full reimbursement."
),
severity="HIGH",
flag_value=claim_amount,
limit_value=policy_limit
)
violations.append(violation)
logger.warning(f"🚩 Rule 2 VIOLATED: claim ${claim_amount} > policy ${policy_limit}")
else:
logger.info(f"✅ Rule 2 PASSED: claim ${claim_amount} ≤ policy ${policy_limit}")
# ── Build the result ───────────────────────────────────
is_suspicious = len(violations) > 0
status = "Suspicious" if is_suspicious else "Valid"
# Build a plain English summary of violations (for the LLM to use!)
if violations:
rule_summary = " | ".join([v.description for v in violations])
else:
rule_summary = (
f"All rules passed. Claim amount ${claim_amount:,.2f} is "
f"within the invoice total ${invoice_total:,.2f} and "
f"policy limit ${policy_limit:,.2f}."
)
result = RuleEngineResult(
is_suspicious=is_suspicious,
status=status,
violations=violations,
claim_amount=claim_amount,
invoice_total=invoice_total,
policy_limit=policy_limit,
rule_summary=rule_summary
)
logger.info(f"📋 Rule Engine verdict: {status} ({len(violations)} violations)")
return result
def to_dict(self, result: RuleEngineResult) -> dict:
"""Convert result to a clean dictionary for JSON response."""
return {
"status": result.status,
"is_suspicious": result.is_suspicious,
"claim_amount": result.claim_amount,
"invoice_total": result.invoice_total,
"policy_limit": result.policy_limit,
"violations": [
{
"rule": v.rule_name,
"description": v.description,
"severity": v.severity
}
for v in result.violations
],
"violation_count": len(result.violations),
"rule_summary": result.rule_summary
}
📋 Step 4 — Key lines explained:
@dataclass→ A Python shortcut for creating simple data containers — like a form with fields!claim_amount > invoice_total→ The core Rule 1 check — pure math, no ambiguity!INVOICE_AMOUNT_UNREADABLE→ If OCR can't read the amount, that's suspicious too. Real fraudsters sometimes submit blurry photos intentionally!rule_summary→ A plain English description we pass to the LLM so it can write a smart explanation
🔹 PHASE 7 — Hugging Face LLM Layer (The Smart Brain! 🧠)
First — Understanding Flan-T5 (Very Important!)
🧒 What is Flan-T5?
Imagine you have a robot friend 🤖 who has read millions of books, articles, and websites. Because it read so much, it learned how to answer questions, explain things, summarize text, and follow instructions — just like a very smart student!
Flan-T5 is exactly that robot. Here are the facts about it:
- 📚 What is it? Flan-T5 is an AI language model made by Google. It was released in 2022 and is available FREE on Hugging Face!
- 📖 What was it trained on? It was trained on over 1,000 different tasks — including question answering, text summarization, translation, and more. Think of it like a student who studied ALL subjects!
- 🎯 Why we use it? We want the model to look at our fraud detection results and write a clear, human-friendly explanation. Flan-T5 is excellent at following instructions and writing explanations.
- 🔡 How it works? It's a "text-to-text" model — you give it text (a prompt/question), it gives you text back (an answer/explanation). Like asking a question and getting an answer!
🚫 What it CANNOT do:
- It cannot access the internet or real-time data
- It cannot see images (that's OCI Vision's job)
- It is NOT 100% accurate — always validate its output with rules first
- The small version (flan-t5-base) has limits on text length
Flan-T5 Size Options
Model Name | Parameters | RAM Needed | Speed | Quality ───────────────────────────────────────────────────────────────────── google/flan-t5-small | 80M | ~300MB | ⚡ Fast | Good google/flan-t5-base | 250M | ~1GB | ✅ Good | Better ← We use this! google/flan-t5-large | 780M | ~3GB | 🐌 Slow | Great google/flan-t5-xl | 3B | ~12GB | 🐌 Slow | Excellent google/flan-t5-xxl | 11B | ~40GB | 🐌 Slow | Best
For our use case, flan-t5-base is the sweet spot — fast enough for a web API, smart enough to write good explanations!
📋 Step 1 — What this code will do:
Load the Flan-T5 model from Hugging Face, build a smart prompt using our invoice data and rule results, and generate a clear human-readable explanation of the fraud verdict.
📋 Step 2 — Why we need it:
Rules tell us WHAT happened ("claim exceeds invoice"). The LLM tells us WHY it matters and explains it in a way a human can easily understand. This is critical for compliance reports and auditor explanations!
# ═══════════════════════════════════════════════════════════════
# llm_engine.py — Hugging Face Flan-T5 Explanation Generator
# ═══════════════════════════════════════════════════════════════
import os
import logging
from typing import Optional
from dotenv import load_dotenv
load_dotenv()
logger = logging.getLogger(__name__)
class LLMEngine:
"""
Uses Hugging Face Flan-T5 to generate human-readable explanations.
🧒 Think of this as a very smart writer ✍️ who takes our fraud
detection results (numbers and rule violations) and writes a
clear, friendly explanation that any person can understand!
It's like having a lawyer who can explain complex rules in simple words!
"""
def __init__(self):
self.model_name = os.getenv("HF_MODEL_NAME", "google/flan-t5-base")
self.model = None
self.tokenizer = None
self._is_loaded = False
logger.info(f"🤖 LLMEngine created — model: {self.model_name}")
logger.info("⏳ Model will be loaded on first use (lazy loading)...")
def _load_model(self):
"""
Load the Flan-T5 model from Hugging Face.
We use "lazy loading" — we only download/load the model when
we first need it. Like not turning on the oven until you're
actually ready to bake! 🍰
First time this runs, it downloads the model (~1GB for base).
After that, it loads from local cache — much faster!
"""
if self._is_loaded:
return # Already loaded, skip!
logger.info(f"📥 Loading Flan-T5 model: {self.model_name}")
logger.info(" (First time may take 1-2 minutes to download...)")
try:
# Import the Hugging Face transformers library
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
import torch
# ── Load the tokenizer ─────────────────────────────
# A tokenizer converts words → numbers the model understands
# Like a translator who converts English to the robot's language!
self.tokenizer = AutoTokenizer.from_pretrained(self.model_name)
# ── Load the actual model ──────────────────────────
# AutoModelForSeq2SeqLM = auto-detect the right model class
# Flan-T5 is a "Seq2Seq" model: sequence-to-sequence
# (input text → output text)
self.model = AutoModelForSeq2SeqLM.from_pretrained(self.model_name)
# ── Use CPU (works on any machine without GPU) ─────
self.device = "cuda" if torch.cuda.is_available() else "cpu"
self.model = self.model.to(self.device)
self.model.eval() # eval() = inference mode (not training)
self._is_loaded = True
logger.info(f"✅ Flan-T5 model loaded successfully on {self.device}!")
except Exception as e:
logger.error(f"❌ Failed to load LLM model: {e}")
raise
def generate_explanation(
self,
invoice_data: dict,
rule_result: dict,
claim_amount: float,
question: Optional[str] = None
) -> str:
"""
Generate a clear explanation of the fraud detection verdict.
This is the core of our LLM layer!
We build a smart prompt → feed it to Flan-T5 → get an explanation back.
Like asking a wise professor: "Look at these facts, what do you think?"
"""
# Load model if not already loaded
self._load_model()
# ── Build the prompt ───────────────────────────────────
# A "prompt" is the text we give to the LLM as input.
# The quality of the prompt directly determines the quality of the answer!
# This is called "prompt engineering" — a key skill in AI! 🎯
prompt = self._build_prompt(invoice_data, rule_result, claim_amount, question)
logger.info(f"🤖 Sending prompt to Flan-T5 ({len(prompt)} chars)...")
try:
# ── Tokenize the prompt ────────────────────────────
# Convert our text prompt → numbers the model understands
# max_length=512 → don't send more than 512 tokens
# truncation=True → cut off if too long
import torch
inputs = self.tokenizer(
prompt,
return_tensors="pt", # "pt" = PyTorch tensors
max_length=512,
truncation=True,
padding=True
).to(self.device)
# ── Generate the model's response ─────────────────
# generate() = "please produce an output for this input"
with torch.no_grad(): # no_grad() = don't update weights (inference only)
output_ids = self.model.generate(
**inputs,
max_new_tokens=200, # Max 200 words in the answer
min_length=30, # At least 30 tokens (avoid one-word answers)
num_beams=4, # Beam search = consider 4 paths, pick best
early_stopping=True, # Stop when a complete sentence is found
no_repeat_ngram_size=3 # Avoid repeating the same 3-word phrases
)
# ── Decode output: numbers → human readable text ───
explanation = self.tokenizer.decode(
output_ids[0],
skip_special_tokens=True # Remove [PAD], [EOS] etc.
)
# Clean up the explanation
explanation = explanation.strip()
if not explanation:
explanation = self._fallback_explanation(rule_result, claim_amount)
logger.info(f"✅ LLM generated explanation ({len(explanation)} chars)")
return explanation
except Exception as e:
logger.error(f"❌ LLM generation failed: {e}")
# Fall back to a rule-based explanation if LLM fails
return self._fallback_explanation(rule_result, claim_amount)
def _build_prompt(
self,
invoice_data: dict,
rule_result: dict,
claim_amount: float,
question: Optional[str] = None
) -> str:
"""
Build a smart, structured prompt for Flan-T5.
Good prompt engineering = good LLM output!
We give the model ALL the context it needs to write a good explanation.
Think of it like giving a student all the facts before asking them
to write a paragraph about it! 📝
"""
vendor_name = invoice_data.get("vendor_name", "Unknown Vendor")
invoice_total = invoice_data.get("invoice_total", "unknown")
invoice_date = invoice_data.get("invoice_date", "unknown")
status = rule_result.get("status", "Unknown")
policy_limit = rule_result.get("policy_limit", "unknown")
rule_summary = rule_result.get("rule_summary", "")
violations = rule_result.get("violations", [])
# Format violations list for the prompt
if violations:
violations_text = "\n".join([
f" - {v['description']}" for v in violations
])
violations_section = f"Rule violations found:\n{violations_text}"
else:
violations_section = "No rule violations found."
# ── The main prompt ────────────────────────────────────
if question:
# If user asked a specific question → answer it!
prompt = f"""You are an invoice compliance expert. Answer this question about an invoice claim.
Invoice Details:
- Vendor: {vendor_name}
- Invoice Total: ${invoice_total}
- Invoice Date: {invoice_date}
- Claim Amount: ${claim_amount}
- Policy Limit: ${policy_limit}
- Status: {status}
- {violations_section}
Question: {question}
Answer clearly and concisely in 2-3 sentences:"""
else:
# Default: generate explanation of the compliance decision
prompt = f"""You are a fraud detection expert. Explain in simple language why this invoice claim is {status}.
Invoice Facts:
- Vendor Name: {vendor_name}
- Invoice Total Amount: ${invoice_total}
- Invoice Date: {invoice_date}
- Amount Being Claimed: ${claim_amount}
- Insurance Policy Limit: ${policy_limit}
- Compliance Status: {status}
- {violations_section}
Provide a clear, professional 2-sentence explanation of this decision:"""
return prompt
def _fallback_explanation(self, rule_result: dict, claim_amount: float) -> str:
"""
Generate a simple rule-based explanation if the LLM fails.
Always have a backup plan! 🔄
"""
rule_summary = rule_result.get("rule_summary", "")
status = rule_result.get("status", "Unknown")
if status == "Suspicious":
return (
f"This claim of ${claim_amount:,.2f} has been flagged as suspicious. "
f"{rule_summary}"
)
else:
return (
f"This claim of ${claim_amount:,.2f} appears to be valid and within "
f"acceptable limits. {rule_summary}"
)
def answer_question(
self,
question: str,
invoice_data: dict,
rule_result: dict,
claim_amount: float
) -> str:
"""
Answer a specific natural language question about the invoice.
Example: "Why was my claim rejected?"
"Is the vendor name matching?"
"When does this invoice expire?"
This makes our API interactive and intelligent! 🗣️
"""
logger.info(f"💬 Answering question: {question}")
return self.generate_explanation(
invoice_data,
rule_result,
claim_amount,
question=question
)
📋 Step 4 — Key lines explained:
AutoTokenizer.from_pretrained("google/flan-t5-base")→ Downloads/loads the tokenizer that converts text to numbers the model understandsAutoModelForSeq2SeqLM→ The correct class for Flan-T5 (it converts input sequence → output sequence)model.eval()→ Puts model in "inference mode" — like telling a student "this is the real exam, no practice now"num_beams=4→ The model explores 4 different answer paths and picks the best one — smarter than just picking the first answer!torch.no_grad()→ Tells PyTorch not to track gradients during inference — saves memory and makes it faster
🔹 PHASE 8 — Decision Engine (Putting It All Together! 🎯)
📋 Step 1 — What this code will do:
This is the conductor 🎵 of our orchestra. It calls OCR, calls Rule Engine, calls LLM — and combines ALL their results into one final clean JSON decision. One function to rule them all!
📋 Step 2 — Why we need it:
Without this, we'd have to wire all three engines together in every endpoint. The Decision Engine creates one single, clean API that our FastAPI layer can call.
# ═══════════════════════════════════════════════════════════════
# decision_engine.py — Combines OCR + Rules + LLM into one verdict
# ═══════════════════════════════════════════════════════════════
import logging
import time
from typing import Optional
from ocr_engine import OCREngine
from rule_engine import RuleEngine
from llm_engine import LLMEngine
from storage_engine import StorageEngine
logger = logging.getLogger(__name__)
class DecisionEngine:
"""
The master coordinator.
Calls OCR → Rule Engine → LLM and returns the final verdict.
🧒 Think of this as the judge in a court case:
- The OCR is the evidence reader (reads the invoice)
- The Rule Engine is the law (checks the rules)
- The LLM is the lawyer (explains the decision in plain language)
- The Decision Engine is the JUDGE who combines everything!
"""
def __init__(self):
# Initialize all sub-engines
self.storage_engine = StorageEngine()
self.ocr_engine = OCREngine()
self.rule_engine = RuleEngine()
self.llm_engine = LLMEngine()
logger.info("✅ DecisionEngine ready — all engines initialized!")
def process_invoice_and_validate(
self,
file_bytes: bytes,
filename: str,
claim_amount: float,
policy_limit: Optional[float] = None
) -> dict:
"""
The complete end-to-end pipeline:
Upload → OCR → Rules → LLM → Decision
This single function replaces having to call each engine separately.
Like pressing ONE button that does ALL the steps! 🚀
"""
start_time = time.time()
logger.info("=" * 60)
logger.info(f"🚀 Starting full invoice validation pipeline")
logger.info(f" File: {filename}, Claim: ${claim_amount:,.2f}")
logger.info("=" * 60)
# ── STEP 1: Upload invoice to OCI Object Storage ──────
logger.info("📤 Step 1/4: Uploading invoice to OCI...")
upload_result = self.storage_engine.upload_invoice(file_bytes, filename)
if not upload_result.get("success"):
return self._error_response(
f"Failed to upload invoice: {upload_result.get('error')}",
claim_amount
)
object_name = upload_result["object_name"]
# ── STEP 2: OCR — Extract invoice data ────────────────
logger.info("🔍 Step 2/4: Extracting invoice data with OCR...")
invoice_data = self.ocr_engine.extract_invoice_data(object_name)
# ── STEP 3: Rule Engine — Check fraud rules ────────────
logger.info("⚖️ Step 3/4: Running fraud detection rules...")
rule_result = self.rule_engine.check(
claim_amount=claim_amount,
invoice_data=invoice_data,
custom_policy=policy_limit
)
rule_dict = self.rule_engine.to_dict(rule_result)
# ── STEP 4: LLM — Generate explanation ────────────────
logger.info("🤖 Step 4/4: Generating LLM explanation...")
explanation = self.llm_engine.generate_explanation(
invoice_data=invoice_data,
rule_result=rule_dict,
claim_amount=claim_amount
)
# ── Build Final Response ───────────────────────────────
processing_time = round(time.time() - start_time, 2)
# THE KEY OUTPUT — exactly what was requested!
final_response = {
# ── PRIMARY OUTPUT (most important!) ──────────────
"status": rule_dict["status"], # "Valid" or "Suspicious"
"reason": explanation, # LLM-generated explanation
# ── Invoice Details ────────────────────────────────
"invoice": {
"vendor_name": invoice_data.get("vendor_name"),
"invoice_total": invoice_data.get("invoice_total"),
"invoice_date": invoice_data.get("invoice_date"),
"invoice_id": invoice_data.get("invoice_id"),
"object_stored": object_name,
},
# ── Claim Details ──────────────────────────────────
"claim": {
"amount_claimed": claim_amount,
"policy_limit": rule_dict.get("policy_limit"),
},
# ── Rule Engine Details ────────────────────────────
"fraud_check": {
"is_suspicious": rule_dict["is_suspicious"],
"violations_found": rule_dict["violation_count"],
"violations": rule_dict["violations"],
},
# ── Meta Information ───────────────────────────────
"meta": {
"processing_time_sec": processing_time,
"service": "Invoice Fraud Detection v1.0",
"llm_model": self.llm_engine.model_name,
}
}
logger.info(f"✅ Pipeline complete! Status: {rule_dict['status']} "
f"| Time: {processing_time}s")
return final_response
def _error_response(self, error_msg: str, claim_amount: float) -> dict:
"""Return a clean error response so the API never crashes."""
return {
"status": "Error",
"reason": f"Processing failed: {error_msg}",
"invoice": {},
"claim": {"amount_claimed": claim_amount},
"fraud_check": {"is_suspicious": None, "violations": []},
"meta": {"error": error_msg}
}
🔹 PHASE 9 — API Layer (The Public Front Door! 🌐)
📋 Step 1 — What this code will do:
Build the full FastAPI web application with three endpoints: upload invoice, validate claim, and ask questions. This is what the outside world talks to!
📋 Step 2 — Why we need it:
All our code so far is "internal". The API layer is what makes it accessible to any user, any frontend, or any other system via HTTP. Like opening the shop to customers!
# ═══════════════════════════════════════════════════════════════
# app.py — Main FastAPI Application
# ═══════════════════════════════════════════════════════════════
# This is the COMPLETE main file. It imports all engines and
# wires them together as HTTP endpoints.
import os
import time
import logging
from typing import Optional
from fastapi import FastAPI, File, UploadFile, HTTPException, Form
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel, Field
import uvicorn
from dotenv import load_dotenv
# ─── Import our engines ───────────────────────────────────────
from decision_engine import DecisionEngine
from llm_engine import LLMEngine
load_dotenv()
# ─── Set up logging ───────────────────────────────────────────
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(name)s — %(message)s"
)
logger = logging.getLogger(__name__)
# ═══════════════════════════════════════════════════════════════
# Initialize FastAPI application
# ═══════════════════════════════════════════════════════════════
app = FastAPI(
title="Invoice Compliance & Fraud Detection API",
description="""
## 🤖 Invoice Fraud Detection System
Uses OCI Document Understanding + Hugging Face Flan-T5 to:
1. Read your invoice (OCR)
2. Check fraud rules
3. Generate AI explanation
### Endpoints:
- **POST /upload-invoice** — Upload invoice + validate claim
- **POST /validate-claim** — Validate with invoice already in OCI storage
- **POST /ask** — Ask a question about an invoice claim
- **GET /health** — Service status check
""",
version="1.0.0",
)
# ── Allow all CORS origins (important for web frontends) ───────
app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_methods=["*"],
allow_headers=["*"]
)
# ── Initialize the Decision Engine once at startup ─────────────
# We initialize it here so it's ready for all requests
decision_engine = DecisionEngine()
llm_engine = LLMEngine() # Separate for /ask endpoint
# ═══════════════════════════════════════════════════════════════
# REQUEST/RESPONSE MODELS (Pydantic)
# These define the exact shape of our API inputs and outputs
# Like a strict form that won't accept wrong data types!
# ═══════════════════════════════════════════════════════════════
class ValidateClaimRequest(BaseModel):
"""Request body for /validate-claim endpoint."""
object_name: str = Field(..., description="OCI Object Storage path to invoice")
claim_amount: float = Field(..., gt=0, description="Amount being claimed ($)")
policy_limit: Optional[float] = Field(None, gt=0, description="Policy limit override ($)")
class AskQuestionRequest(BaseModel):
"""Request body for /ask endpoint."""
question: str = Field(..., description="Your natural language question")
claim_amount: float = Field(..., description="Amount being claimed ($)")
invoice_total: Optional[float] = Field(None, description="Invoice total amount ($)")
vendor_name: Optional[str] = Field(None, description="Vendor name")
invoice_date: Optional[str] = Field(None, description="Invoice date")
status: Optional[str] = Field(None, description="Fraud detection status")
# ═══════════════════════════════════════════════════════════════
# ENDPOINT 1: GET /health — Is our service alive? 🟢
# ═══════════════════════════════════════════════════════════════
@app.get("/health", summary="Health Check")
async def health_check():
"""Simple health check — returns status of the service."""
return {
"status": "healthy",
"service": "Invoice Fraud Detection API",
"version": "1.0.0",
"timestamp": time.time(),
"engines": {
"decision_engine": "ready",
"llm_model": os.getenv("HF_MODEL_NAME", "google/flan-t5-base"),
"policy_limit": f"${os.getenv('POLICY_LIMIT', '5000.0')}"
}
}
# ═══════════════════════════════════════════════════════════════
# ENDPOINT 2: POST /upload-invoice
# The MAIN endpoint! Upload invoice image + submit claim → get verdict
# ═══════════════════════════════════════════════════════════════
@app.post(
"/upload-invoice",
summary="Upload Invoice & Validate Claim",
description="""
Upload an invoice image/PDF and submit a claim amount.
The system will:
1. Store the invoice in OCI Object Storage
2. Extract text using OCI Document Understanding
3. Check fraud rules
4. Generate AI explanation using Flan-T5
Returns: {"status": "Valid/Suspicious", "reason": "explanation..."}
"""
)
async def upload_invoice_and_validate(
file: UploadFile = File(..., description="Invoice file (JPEG, PNG, PDF)"),
claim_amount: float = Form(..., description="Amount being claimed in dollars"),
policy_limit: Optional[float] = Form(None, description="Optional policy limit override")
):
"""
Complete pipeline:
Upload invoice → OCR → Rules → LLM → Return verdict JSON
"""
logger.info(f"📨 POST /upload-invoice received")
logger.info(f" File: {file.filename}, Claim: ${claim_amount}")
# ── Validate file type ─────────────────────────────────────
allowed_types = {
"image/jpeg", "image/jpg", "image/png",
"image/tiff", "application/pdf"
}
if file.content_type not in allowed_types:
raise HTTPException(
status_code=400,
detail={
"error": "Unsupported file type",
"received": file.content_type,
"allowed": list(allowed_types)
}
)
# ── Read the file bytes ────────────────────────────────────
try:
file_bytes = await file.read()
except Exception as e:
raise HTTPException(status_code=500, detail=f"File read error: {e}")
# ── File size check (20MB max) ─────────────────────────────
if len(file_bytes) > 20 * 1024 * 1024:
raise HTTPException(
status_code=413,
detail="File too large. Maximum size: 20MB"
)
# ── Validate claim amount ──────────────────────────────────
if claim_amount <= 0:
raise HTTPException(
status_code=400,
detail="claim_amount must be greater than 0"
)
# ── Run the complete pipeline ──────────────────────────────
try:
result = decision_engine.process_invoice_and_validate(
file_bytes=file_bytes,
filename=file.filename,
claim_amount=claim_amount,
policy_limit=policy_limit
)
return JSONResponse(content=result, status_code=200)
except Exception as e:
logger.error(f"❌ Pipeline error: {e}")
raise HTTPException(status_code=500, detail=f"Processing error: {str(e)}")
# ═══════════════════════════════════════════════════════════════
# ENDPOINT 3: POST /validate-claim
# Validate against an invoice already in OCI Object Storage
# ═══════════════════════════════════════════════════════════════
@app.post(
"/validate-claim",
summary="Validate Claim Against Stored Invoice"
)
async def validate_claim(request: ValidateClaimRequest):
"""
Validate a claim amount against an invoice already in OCI Object Storage.
Useful when the invoice was uploaded in a previous step.
"""
logger.info(f"📨 POST /validate-claim: {request.object_name}, ${request.claim_amount}")
try:
# OCR the stored invoice
from ocr_engine import OCREngine
from rule_engine import RuleEngine
ocr = OCREngine()
rules = RuleEngine()
invoice_data = ocr.extract_invoice_data(request.object_name)
rule_result = rules.check(
claim_amount=request.claim_amount,
invoice_data=invoice_data,
custom_policy=request.policy_limit
)
rule_dict = rules.to_dict(rule_result)
explanation = llm_engine.generate_explanation(
invoice_data=invoice_data,
rule_result=rule_dict,
claim_amount=request.claim_amount
)
return JSONResponse(content={
"status": rule_dict["status"],
"reason": explanation,
"invoice": invoice_data,
"fraud_check": rule_dict
})
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
# ═══════════════════════════════════════════════════════════════
# ENDPOINT 4: POST /ask (Optional — Natural Language Q&A!)
# Ask a question about an invoice in plain English
# ═══════════════════════════════════════════════════════════════
@app.post(
"/ask",
summary="Ask Natural Language Question About Invoice",
description="""
Ask any question about an invoice claim in plain English.
Example questions:
- "Why was my claim rejected?"
- "Is the claim amount within the policy limit?"
- "What does the fraud detection result mean?"
"""
)
async def ask_question(request: AskQuestionRequest):
"""
Natural language Q&A about an invoice claim.
Uses Flan-T5 to answer any question about the invoice data!
"""
logger.info(f"💬 POST /ask: {request.question}")
# Build simplified invoice and rule data from request
invoice_data = {
"vendor_name": request.vendor_name or "Unknown",
"invoice_total": request.invoice_total,
"invoice_date": request.invoice_date or "Unknown",
}
# Build a minimal rule result
is_suspicious = request.status == "Suspicious" if request.status else False
rule_result = {
"status": request.status or "Unknown",
"is_suspicious": is_suspicious,
"policy_limit": os.getenv("POLICY_LIMIT", "5000.0"),
"violations": [],
"rule_summary": f"Claim of ${request.claim_amount} reviewed."
}
try:
answer = llm_engine.answer_question(
question=request.question,
invoice_data=invoice_data,
rule_result=rule_result,
claim_amount=request.claim_amount
)
return JSONResponse(content={
"question": request.question,
"answer": answer,
"model": llm_engine.model_name
})
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
# ═══════════════════════════════════════════════════════════════
# Start the server
# ═══════════════════════════════════════════════════════════════
if __name__ == "__main__":
port = int(os.getenv("APP_PORT", 8000))
print("\n" + "=" * 65)
print("🚀 Invoice Compliance & Fraud Detection API Starting...")
print("=" * 65)
print(f" API URL: http://localhost:{port}")
print(f" Swagger Docs: http://localhost:{port}/docs")
print(f" Health Check: http://localhost:{port}/health")
print(f" Policy Limit: ${os.getenv('POLICY_LIMIT', '5000.00')}")
print(f" LLM Model: {os.getenv('HF_MODEL_NAME', 'google/flan-t5-base')}")
print("=" * 65 + "\n")
uvicorn.run("app:app", host="0.0.0.0", port=port, reload=False)
📋 Step 4 — Key lines explained:
@app.post("/upload-invoice")→ Creates the endpoint at this URL. ThePOSTmethod is for sending data to the server.UploadFile = File(...)→ Accepts a file upload....means required.claim_amount: float = Form(...)→ Accepts a form field alongside the file uploadJSONResponse(content=result)→ Returns our result as a JSON HTTP response
🔹 PHASE 10 — requirements.txt (The Shopping List 🛒)
📋 Step 1 — What this does: Lists every Python library our system needs. One command installs everything!
# requirements.txt # ───────────────────────────────────────────────────────────── # Install everything with: pip install -r requirements.txt # ───────────────────────────────────────────────────────────── # ── Oracle Cloud Infrastructure SDK ─────────────────────────── # Connects to OCI Object Storage + Document Understanding oci>=2.120.0 # ── FastAPI + Web Server ─────────────────────────────────────── fastapi>=0.110.0 uvicorn[standard]>=0.29.0 python-multipart>=0.0.9 # For file uploads # ── Settings / Environment ───────────────────────────────────── python-dotenv>=1.0.0 pydantic>=2.6.0 # ── Hugging Face — LLM (Flan-T5) ────────────────────────────── transformers>=4.40.0 # Core HuggingFace library torch>=2.0.0 # PyTorch (model runs on this) sentencepiece>=0.1.99 # Tokenizer required by Flan-T5 accelerate>=0.29.0 # Makes model loading faster # ── Image Processing ─────────────────────────────────────────── Pillow>=10.3.0 # Image validation # ── HTTP Requests ───────────────────────────────────────────── requests>=2.31.0 # ── Testing ─────────────────────────────────────────────────── pytest>=8.0.0 httpx>=0.27.0 # For testing FastAPI endpoints
🔹 PHASE 11 — Dockerfile (Packing the Box! 📦)
📋 Step 1 — What this does:
Creates a Docker container image with ALL our code, dependencies, and settings baked in. Like packing your entire kitchen into a portable box — and it works the same wherever you open it!
# ═══════════════════════════════════════════════════════════════
# Dockerfile — Recipe for packaging our Fraud Detection System
# ═══════════════════════════════════════════════════════════════
# ─── Start from official Python 3.11 image ────────────────────
FROM python:3.11-slim-bookworm
# ─── Labels ───────────────────────────────────────────────────
LABEL maintainer="your-email@example.com"
LABEL description="Invoice Fraud Detection System — OCI + Flan-T5"
LABEL version="1.0.0"
# ─── Install system dependencies ──────────────────────────────
# These are needed for PyTorch and image processing libraries
RUN apt-get update && apt-get install -y \
gcc \
g++ \
libpq-dev \
libjpeg-dev \
libpng-dev \
curl \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
# ─── Set working directory ─────────────────────────────────────
WORKDIR /app
# ─── Install Python dependencies FIRST ────────────────────────
# We do this before copying code for Docker layer caching.
# If requirements.txt hasn't changed, Docker skips this step on rebuild!
COPY requirements.txt .
RUN pip install --no-cache-dir --upgrade pip \
&& pip install --no-cache-dir -r requirements.txt
# ─── Pre-download the Flan-T5 model ───────────────────────────
# Download model at BUILD time (not runtime) so startup is fast!
# Like packing the brain into the box before shipping it.
# Without this, the first request would take 2+ minutes to download.
ARG HF_MODEL=google/flan-t5-base
RUN python -c "from transformers import AutoTokenizer, AutoModelForSeq2SeqLM; \
print('Downloading Flan-T5 model...'); \
AutoTokenizer.from_pretrained('${HF_MODEL}'); \
AutoModelForSeq2SeqLM.from_pretrained('${HF_MODEL}'); \
print('Model downloaded and cached!')"
# ─── Copy application code ─────────────────────────────────────
COPY app.py .
COPY decision_engine.py .
COPY ocr_engine.py .
COPY rule_engine.py .
COPY llm_engine.py .
COPY storage_engine.py .
COPY .env .
COPY oci_api_key.pem .
# ─── Security — run as non-root user ──────────────────────────
RUN useradd --create-home --shell /bin/bash appuser \
&& chown -R appuser:appuser /app
USER appuser
# ─── Expose port ──────────────────────────────────────────────
EXPOSE 8000
# ─── Health check ─────────────────────────────────────────────
HEALTHCHECK --interval=30s --timeout=15s --start-period=30s --retries=3 \
CMD curl -f http://localhost:8000/health || exit 1
# ─── Start command ────────────────────────────────────────────
CMD ["python", "app.py"]
Build and Run Docker Commands
# ── Step 1: Build the image ──────────────────────────────────── # Note: First build takes 5-10 mins (downloads PyTorch + Flan-T5!) docker build -t invoice-fraud-system:v1.0 . # ── Step 2: Run locally to test ─────────────────────────────── docker run -d \ --name fraud-detector \ -p 8000:8000 \ invoice-fraud-system:v1.0 # ── Step 3: Check it's running ──────────────────────────────── docker logs fraud-detector curl http://localhost:8000/health # ── Step 4: Push to OCI Container Registry ──────────────────── docker tag invoice-fraud-system:v1.0 \ iad.ocir.io/<your-namespace>/fraud-detection/invoice-fraud:v1.0 docker push iad.ocir.io/<your-namespace>/fraud-detection/invoice-fraud:v1.0
🔹 PHASE 12 — Deployment on OCI
Option A: Deploy on OCI Compute VM (Simplest!)
# ── On your OCI Compute VM (SSH in first) ───────────────────── # Step 1: Install Docker on the VM sudo apt-get update sudo apt-get install -y docker.io sudo systemctl start docker sudo systemctl enable docker sudo usermod -aG docker $USER # Step 2: Login to OCI Container Registry docker login iad.ocir.io \ -u <namespace>/<username> \ -p <your-auth-token> # Step 3: Pull and run your container docker pull iad.ocir.io/<namespace>/fraud-detection/invoice-fraud:v1.0 docker run -d \ --name fraud-detector \ -p 80:8000 \ --restart always \ iad.ocir.io/<namespace>/fraud-detection/invoice-fraud:v1.0 # Step 4: Open port 80 in OCI Security List # (OCI Console → VCN → Security Lists → Add Ingress Rule → Port 80) # Step 5: Your API is now public! # Access it at: http://<your-vm-public-ip>/docs
Option B: Deploy on OCI Kubernetes (OKE)
# kubernetes-deployment.yaml
# Deploy our fraud detection system to OKE
apiVersion: apps/v1
kind: Deployment
metadata:
name: fraud-detector
labels:
app: fraud-detector
spec:
replicas: 2 # Run 2 copies for reliability
selector:
matchLabels:
app: fraud-detector
template:
metadata:
labels:
app: fraud-detector
spec:
containers:
- name: fraud-detector
image: iad.ocir.io/<namespace>/fraud-detection/invoice-fraud:v1.0
ports:
- containerPort: 8000
resources:
requests:
memory: "2Gi" # Flan-T5 needs at least 1GB RAM
cpu: "500m"
limits:
memory: "4Gi"
cpu: "2000m"
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 30
periodSeconds: 30
imagePullSecrets:
- name: ocir-secret # Kubernetes secret for OCIR login
---
apiVersion: v1
kind: Service
metadata:
name: fraud-detector-service
spec:
selector:
app: fraud-detector
ports:
- port: 80
targetPort: 8000
type: LoadBalancer # Creates a public IP automatically!
# Apply the deployment to OKE kubectl apply -f kubernetes-deployment.yaml # Check pods are running kubectl get pods # Get the public IP of our service kubectl get service fraud-detector-service # Look for EXTERNAL-IP column — that's your public URL!
🔹 PHASE 13 — Testing with Real Examples 🧪
Test 1: Valid Claim (Should Pass ✅)
Scenario: Invoice is $500, claiming $400, policy limit is $5000. All good!
# Test with a real invoice file curl -X POST http://localhost:8000/upload-invoice \ -F "file=@sample_invoice.jpg" \ -F "claim_amount=400.00"
Expected output:
{
"status": "Valid",
"reason": "The claim of $400.00 is within the invoice total of $500.00 and
the policy limit of $5000.00. This claim appears legitimate
and can be processed for reimbursement.",
"invoice": {
"vendor_name": "ACME Corp",
"invoice_total": 500.00,
"invoice_date": "2026-04-01",
"invoice_id": "INV-2026-001",
"object_stored": "invoices/a1b2c3d4_sample_invoice.jpg"
},
"claim": {
"amount_claimed": 400.00,
"policy_limit": 5000.00
},
"fraud_check": {
"is_suspicious": false,
"violations_found": 0,
"violations": []
},
"meta": {
"processing_time_sec": 4.73,
"service": "Invoice Fraud Detection v1.0",
"llm_model": "google/flan-t5-base"
}
}
Test 2: Suspicious Claim — Exceeds Invoice 🚩
Scenario: Invoice is $500, claiming $800. Fraud detected!
curl -X POST http://localhost:8000/upload-invoice \ -F "file=@sample_invoice.jpg" \ -F "claim_amount=800.00"
Expected output:
{
"status": "Suspicious",
"reason": "The claim amount of $800.00 exceeds the invoice total of $500.00 by $300.00.
This claim has been flagged because it is not possible to be reimbursed
for more than the actual invoice amount paid.",
"invoice": {
"vendor_name": "ACME Corp",
"invoice_total": 500.00,
"invoice_date": "2026-04-01",
"invoice_id": "INV-2026-001"
},
"claim": {
"amount_claimed": 800.00,
"policy_limit": 5000.00
},
"fraud_check": {
"is_suspicious": true,
"violations_found": 1,
"violations": [
{
"rule": "CLAIM_EXCEEDS_INVOICE",
"description": "The claim amount $800.00 exceeds the invoice total $500.00 by
$300.00. A claim cannot be larger than what was actually paid.",
"severity": "HIGH"
}
]
}
}
Test 3: Ask a Question (Natural Language!) 💬
curl -X POST http://localhost:8000/ask \
-H "Content-Type: application/json" \
-d '{
"question": "Why was my claim rejected?",
"claim_amount": 800.00,
"invoice_total": 500.00,
"vendor_name": "ACME Corp",
"status": "Suspicious"
}'
Expected output:
{
"question": "Why was my claim rejected?",
"answer": "Your claim of $800.00 was rejected because it exceeds the invoice total
of $500.00. Insurance policies only reimburse for amounts actually paid,
so a claim cannot be higher than the invoice.",
"model": "google/flan-t5-base"
}
Test with Python Script
# test_system.py — Test all three scenarios
import requests
import json
BASE_URL = "http://localhost:8000"
def test_health():
"""Test if the service is running."""
r = requests.get(f"{BASE_URL}/health")
print("✅ Health:", r.json()["status"])
def test_valid_claim(invoice_path, claim_amount):
"""Test a valid claim."""
with open(invoice_path, "rb") as f:
r = requests.post(
f"{BASE_URL}/upload-invoice",
files={"file": f},
data={"claim_amount": claim_amount}
)
result = r.json()
print(f"\n{'='*50}")
print(f"Claim: ${claim_amount}")
print(f"Status: {result['status']}")
print(f"Reason: {result['reason']}")
print(f"Violations: {result['fraud_check']['violations_found']}")
def test_ask_question(question, claim_amount, invoice_total):
"""Test natural language Q&A."""
r = requests.post(
f"{BASE_URL}/ask",
json={
"question": question,
"claim_amount": claim_amount,
"invoice_total": invoice_total,
"status": "Suspicious" if claim_amount > invoice_total else "Valid"
}
)
result = r.json()
print(f"\nQ: {result['question']}")
print(f"A: {result['answer']}")
# Run all tests!
test_health()
test_valid_claim("sample_invoice.jpg", 400.00) # Should be Valid
test_valid_claim("sample_invoice.jpg", 800.00) # Should be Suspicious
test_ask_question("Why was my claim rejected?", 800.00, 500.00)
🔹 PHASE 14 — Final Project Structure & Summary
Complete Project Structure
invoice-fraud-detection/ │ ├── 📄 app.py ← FastAPI application (main entry point) ├── 📄 decision_engine.py ← Master coordinator (orchestrates all engines) ├── 📄 storage_engine.py ← OCI Object Storage upload logic ├── 📄 ocr_engine.py ← OCI Document Understanding OCR ├── 📄 rule_engine.py ← Fraud detection rules ├── 📄 llm_engine.py ← Hugging Face Flan-T5 LLM │ ├── 📄 requirements.txt ← All Python dependencies ├── 📄 Dockerfile ← Container packaging ├── 📄 kubernetes-deployment.yaml ← K8s deployment (optional) │ ├── 📄 .env ← Secret settings ← NEVER commit to Git! ├── 📄 .gitignore ← Protects secrets ├── 🔑 oci_api_key.pem ← OCI private key ← NEVER commit to Git! │ └── 📄 test_system.py ← Test all endpoints
What We Built — The Complete Summary
- storage_engine.py → Uploads invoices to Oracle Cloud Object Storage. Every invoice gets a unique ID so nothing gets overwritten.
- ocr_engine.py → Uses OCI Document Understanding (pre-trained on invoices!) to extract vendor name, invoice total, and date automatically. No manual reading!
- rule_engine.py → A strict detective that checks two rules: claim must not exceed invoice, and claim must not exceed policy limit. Returns violations in plain English.
- llm_engine.py → Loads Google's Flan-T5 model from Hugging Face. Takes all the facts and generates a clear, human-readable explanation. Can also answer questions!
- decision_engine.py → The master coordinator that calls all engines in the right order and bundles everything into one clean response.
- app.py → FastAPI web server with 4 endpoints: health check, upload invoice, validate claim, and ask questions. Has auto-generated Swagger docs at
/docs! - Dockerfile → Packages everything including the pre-downloaded Flan-T5 model into a portable container that runs anywhere.
Technologies You Used
- ☁️ OCI Object Storage — Cloud file storage for invoices
- 🤖 OCI Document Understanding — AI-powered invoice OCR (pre-trained for invoices)
- 🧠 Hugging Face Flan-T5 — Open-source LLM for generating explanations
- 🌐 FastAPI — Modern Python web framework for building APIs
- 🐳 Docker — Container packaging for deployment anywhere
- ⚓ OKE / OCI Compute — Cloud deployment options
Happy coding! 🤖✨
Comments
Post a Comment