Skip to main content

OCI Vision Service: Image Analysis, OCR, and Computer Vision

Calculating read time…

Imagine you run a factory that makes biscuits.  Every biscuit on the conveyor belt must be the perfect shape and colour. A cracked or burnt biscuit must be caught and removed before it reaches the customer. With 10,000 biscuits per hour, no human can possibly check every single one.

What if a camera could watch the belt and instantly flag every broken biscuit — automatically? That's exactly the kind of problem OCI Vision solves! It gives your applications the ability to look at images and understand what they see — just like how a human expert would, but at machine speed and scale.




What Is OCI Vision?

OCI Vision is a cloud AI service that can look at a photo or video and understand what's in it. This ability — teaching computers to understand images — is called Computer Vision.

Think of it like training a very attentive robot to look at pictures. You show it millions of photos during training. It learns patterns. Then when you show it a new photo it has never seen before, it can tell you: "This is a cat", "There is a car on the left side", or "This text says OPEN".

💡 Real-world analogy: Think of OCI Vision like a super-smart intern at your company. You give them a photo and they can tell you everything in it: what objects are visible, where each object is, what text is written, and how many people are in the image — all in under one second!

What Can OCI Vision Actually Do? 

OCI Vision has two main capability groups — think of them as two different types of "eyes":

  • 🖼️ Image Analysis — Looks at photographs and scenes to understand the content
  • 📄 Document AI — Looks at documents (receipts, forms, invoices) to extract structured data

Under these two groups, here are all the specific features available:

  • 🏷️ Image Classification — "What is the overall subject of this photo?" (e.g. Beach, Forest, Stadium)
  • 📦 Object Detection — "What specific objects are in this photo, and where exactly are they?" (returns bounding box coordinates)
  • 👤 Face Detection — "Are there human faces in this image, and where?" (returns bounding boxes around each face)
  • 🔤 Text Detection (OCR) — "Is there any text written in this image? Read it for me."
  • 🎬 Video Analysis — Analyse every frame of a stored video for objects, labels, text, and faces
  • 📡 Live Stream Analysis — Real-time AI analysis of live video streams
  • 🏗️ Custom Model Training — Train your own vision model on YOUR specific images and labels

Let's understand each one deeply, starting from the simplest! 🎯

The Big Picture — How OCI Vision Works 🗺️

YOUR IMAGE (jpg / png / bmp / tiff / gif)
        │
        │  Two ways to send the image:
        │  Option A: Base64 encode it (embed it directly in the API call)
        │  Option B: Store it in OCI Object Storage and point to it
        │
        ▼
┌─────────────────────────────────────────────────────────────┐
│                   OCI Vision Service                        │
│                                                             │
│  You tell it WHAT to look for (the "features" list):        │
│  ┌──────────────────────────────────────────────────────┐   │
│  │ IMAGE_CLASSIFICATION  — What is in the whole image?  │   │
│  │ OBJECT_DETECTION      — What objects? Where exactly? │   │
│  │ TEXT_DETECTION        — What text is visible?        │   │
│  │ FACE_DETECTION        — Are there faces? Where?      │   │
│  └──────────────────────────────────────────────────────┘   │
│                                                             │
│  Deep Learning models analyse the image                    │
│  Results returned as structured JSON                       │
└─────────────────────────────────────────────────────────────┘
        │
        │  Returns structured JSON with results + confidence scores
        ▼
YOUR APPLICATION 🎉
(shows labels, draws bounding boxes, extracts text, etc.)

You can ask OCI Vision to do multiple things at the same time in a single API call. For example: "In this image, please find all objects AND read any visible text AND detect faces — all in one go." That saves API calls and is much faster!

Feature 1: Image Classification — What's in This Photo? 🏷️

What Is Image Classification?

Image Classification answers the question: "What is the main subject or theme of this entire image?"

Imagine you are sorting a pile of 10,000 photos from your holiday. You want to put all beach photos in one folder, all mountain photos in another, all city photos in a third. Doing that by hand would take days! OCI Vision's Image Classification can sort all 10,000 photos in minutes — automatically.

Example — you send this photo:
[A photo of a football match in a stadium]

OCI Vision Image Classification returns:
┌─────────────────────────────────────┐
│  Labels (what it thinks the image   │
│  is about, from most to least sure):│
│                                     │
│  Sports          confidence: 0.98   │
│  Stadium         confidence: 0.97   │
│  Football        confidence: 0.93   │
│  Crowd           confidence: 0.91   │
│  Grass           confidence: 0.85   │
│  Night           confidence: 0.79   │
└─────────────────────────────────────┘

Each label comes with a confidence score between 0 and 1. A score of 0.98 means the AI is 98% confident that label is correct. You can set a threshold — for example, only use labels above 0.80.

✅ Great use cases for Image Classification:
📸 Automatically organise thousands of product photos by category
🏭 Classify factory floor images as "normal" or "needs inspection"
📰 Tag news article images automatically (Sports, Politics, Nature, etc.)
🛒 E-commerce product cataloguing — sort uploaded images by type automatically

Feature 2: Object Detection — Find Objects AND Their Location 📦

What Is Object Detection?

Image Classification tells you the theme of the whole photo. Object Detection goes much further — it finds every individual object AND tells you exactly where in the image each object is located.

The location is described using a bounding box — four numbers that define a rectangle around the object: the top-left corner's X and Y coordinates, plus the width and height of the rectangle. Think of it like drawing a rectangle around each object with a ruler!

Example — you send this photo:
[A busy street with 2 cars, 3 people, and a traffic light]

OCI Vision Object Detection returns:
┌──────────────────────────────────────────────────────────────┐
│  Object: "Car"         Confidence: 0.96                      │
│  Bounding Box: x=45, y=120, width=200, height=95            │
│  (= there's a car starting at pixel 45,120 — 200px wide)    │
├──────────────────────────────────────────────────────────────┤
│  Object: "Car"         Confidence: 0.94                      │
│  Bounding Box: x=320, y=115, width=185, height=90           │
├──────────────────────────────────────────────────────────────┤
│  Object: "Person"      Confidence: 0.98                      │
│  Bounding Box: x=10, y=80, width=50, height=140             │
├──────────────────────────────────────────────────────────────┤
│  Object: "Person"      Confidence: 0.97                      │
│  Bounding Box: x=70, y=85, width=48, height=135             │
├──────────────────────────────────────────────────────────────┤
│  Object: "Person"      Confidence: 0.89                      │
│  Bounding Box: x=550, y=90, width=52, height=130            │
├──────────────────────────────────────────────────────────────┤
│  Object: "Traffic Light" Confidence: 0.92                   │
│  Bounding Box: x=430, y=20, width=30, height=80             │
└──────────────────────────────────────────────────────────────┘

With these bounding box coordinates, your application can draw coloured rectangles on the image — the classic "boxes around detected objects" you've seen in AI demos. It can also count objects ("there are 2 cars and 3 people in this scene") or trigger alerts ("a person has entered the restricted zone!").

✅ Great use cases for Object Detection:
🚗 Counting vehicles in a car park from a security camera image
📦 Counting packages on a warehouse conveyor belt
🏭 Detecting defective items on a production line
🌿 Detecting if vegetation is growing too close to power lines (safety monitoring)
🛒 Counting products on a retail shelf to trigger restocking alerts

Feature 3: Face Detection — Find Every Face in a Photo 👤

What Is Face Detection?

Face Detection is a specialised version of Object Detection — but focused entirely on human faces. OCI Vision finds every face in an image and returns a bounding box for each one.

It also detects facial landmarks — key points on the face like the position of the eyes, nose, and mouth corners. These coordinates are useful for applications that need to measure facial geometry (like blurring faces for privacy, or aligning faces for comparison).

Example — you send a team photo with 6 people:

OCI Vision Face Detection returns:
┌────────────────────────────────────────────────────────────┐
│  Face 1: Bounding Box: x=50, y=30, width=80, height=90    │
│  Face 2: Bounding Box: x=180, y=25, width=78, height=88   │
│  Face 3: Bounding Box: x=310, y=35, width=82, height=91   │
│  Face 4: Bounding Box: x=440, y=28, width=76, height=87   │
│  Face 5: Bounding Box: x=570, y=32, width=79, height=89   │
│  Face 6: Bounding Box: x=700, y=30, width=81, height=90   │
│                                                            │
│  (6 faces found in total)                                 │
└────────────────────────────────────────────────────────────┘
✅ Great use cases for Face Detection:
🔏 Privacy compliance: Automatically blur faces in photos before publishing
📸 Media management: Find all photos where faces appear for tagging
🎓 Attendance: Detect how many people attended an event from photos
🛡️ Security: Detect presence of people in restricted areas from camera feeds
❌ Important: OCI Vision Face Detection identifies WHERE faces are (detection + bounding boxes). It does NOT perform facial recognition — it will NOT tell you WHO the person is by name. OCI Vision does not have a "face recognition / identity matching" feature. Always respect privacy laws when using face detection in your applications!

Feature 4: Text Detection (OCR) — Read Text in Images 🔤

What Is Text Detection?

OCR stands for Optical Character Recognition. It means reading text that appears visually in an image — like a sign on a wall, text printed on a package, a handwritten note, or a number plate on a car.

Think of it like having a friend who can look at any photo and type out every single word they see — whether it's on a shop sign, a whiteboard, or even a photo of a handwritten letter.

Example — you send a photo of a street scene with a billboard:

OCI Vision Text Detection returns:
┌──────────────────────────────────────────────────────────────┐
│  Text Block 1:                                               │
│  Text: "GRAND OPENING"                                       │
│  Bounding Box: x=100, y=50, width=350, height=60           │
│  Confidence: 0.99                                            │
│                                                              │
│  Text Block 2:                                               │
│  Text: "Saturday 15th March"                                 │
│  Bounding Box: x=120, y=120, width=310, height=45          │
│  Confidence: 0.97                                            │
│                                                              │
│  Text Block 3:                                               │
│  Text: "FREE ENTRY"                                          │
│  Bounding Box: x=150, y=175, width=260, height=55          │
│  Confidence: 0.98                                            │
└──────────────────────────────────────────────────────────────┘

Each detected text block comes with its text content, its location (bounding box), and a confidence score. OCI Vision Text Detection can read printed text, handwriting, text in multiple languages, and even text at unusual angles.

✅ Great use cases for Text Detection:
🚗 Reading number plates from parking lot camera images
📦 Scanning product labels and barcodes in warehouse photos
📋 Reading handwritten form responses from scanned paper forms
🌐 Extracting text from website screenshots for accessibility tools
🏷️ Reading price tags from retail shelf photos automatically

Feature 5: Video Analysis — AI Eyes on Every Frame 🎬

What Is Video Analysis?

A video is just thousands of photos (called frames) played very quickly. OCI Vision's Video Analysis applies its AI models to every single frame of a stored video file and returns what it found at each timestamp.

You upload your video to OCI Object Storage, submit it for analysis, and OCI Vision processes every frame. The output tells you exactly when in the video each object, label, text, or face appeared.

Example — you submit a 2-minute security camera recording:

OCI Vision Video Analysis output (simplified):
┌─────────────────────────────────────────────────────────────┐
│  Timestamp: 0:00–0:15  │ Objects: Car, Road                 │
│  Timestamp: 0:16–0:22  │ Objects: Car, Person               │
│  Timestamp: 0:23–0:23  │ Objects: Person ← entered zone!    │
│  Timestamp: 0:24–0:45  │ Objects: Car, Person, Bicycle      │
│  Timestamp: 0:46–1:02  │ Objects: Car, Road                 │
│  Timestamp: 1:03–1:03  │ Text: "Exit 4" detected            │
│  Timestamp: 1:04–2:00  │ Objects: Car, Road                 │
└─────────────────────────────────────────────────────────────┘

Use the timeline to jump directly to 0:23 to review the person
who entered the restricted zone! 🚨
💡 The Timeline Bar: In the OCI Console, Video Analysis results include an interactive timeline bar. You can click on any label or object and it jumps you to the exact frame in the video where it appears. No more manually scrubbing through hours of security footage!
✅ Great use cases for Video Analysis:
🔒 Security: Review hours of CCTV footage automatically — jump directly to frames with people
🏭 Manufacturing: Analyse production line video to detect defects at specific timestamps
📺 Media: Index video content automatically so viewers can search by what appears on screen
🚦 Traffic: Analyse road camera footage to count vehicles and detect incidents

Setting Up Your Environment — The One-Time Setup 🛠️

Before we write any Python code, we need to set up our OCI credentials. This is a one-time process — once done, all your future OCI code will just work!

Step 1: Install the OCI Python SDK

📌 What this command does:
This installs the official Oracle Python package onto your computer. It's like downloading an app — once installed, you can use all OCI services from Python code. The oci package contains everything you need to talk to OCI Vision.
pip install oci

Step 2: Create Your OCI API Key

  • Log in to OCI Console → click your user avatar (top-right) → My Profile
  • Scroll down to API Keys → click Add API Key
  • Choose Generate API Key Pair → click Download Private Key
  • Save the private key file as ~/.oci/oci_api_key.pem
  • Copy the configuration snippet shown on screen — you'll need it for the next step

Step 3: Create the OCI Config File

📌 What this file does:
The config file is your ID card for OCI. It tells the Python SDK who you are and which account you belong to. Without it, every API call will return "Unauthorized". Create a file at ~/.oci/config and paste the values you copied from the console.
[DEFAULT]
user=ocid1.user.oc1..aaaaaaaa...your-user-ocid...
fingerprint=ab:cd:ef:12:34:56:78:90:ab:cd:ef:12:34:56:78:90
tenancy=ocid1.tenancy.oc1..aaaaaaaa...your-tenancy-ocid...
region=us-ashburn-1
key_file=~/.oci/oci_api_key.pem

Step 4: Add the IAM Policy (Admin Step)

📌 What this does:
OCI requires an administrator to explicitly grant permission to use each service. An admin in your account needs to add this policy statement. Think of it like your company's IT department granting you access to a new software system. Without this, every API call fails with "Authorization failed".
allow group <your-group-name> to use ai-service-vision-family in tenancy

Your First Python Code — Analyse a Local Image! 🐍

Part 1 — Image Classification: What Is in This Photo?

📌 What the code below does — in plain English:
This code does 4 things:
  1. Reads an image file from your computer and converts it to Base64 format (a special text encoding that APIs can safely transmit)
  2. Connects to OCI Vision using your config file
  3. Sends the image and asks: "Please classify this image — what is it about?"
  4. Prints all the labels it found, sorted from most confident to least confident
This is equivalent to showing a photo to your AI intern and asking "What do you see overall?"
import oci
import base64

# ── STEP 1: Load OCI credentials from config file ────────────────────────────
config = oci.config.from_file()

# ── STEP 2: Connect to OCI Vision service ────────────────────────────────────
# This opens a connection to OCI Vision — like dialling the service's phone number
ai_vision_client = oci.ai_vision.AIServiceVisionClient(config)

# ── STEP 3: Your compartment ID ──────────────────────────────────────────────
# Find this in OCI Console → Identity → Compartments
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."

# ── STEP 4: Convert your image file to Base64 ────────────────────────────────
# APIs cannot send raw binary image files directly.
# Base64 converts the image into a text string that can be safely included in the API request.
# Think of it like translating the image into a language the API can read!
def encode_image_to_base64(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

# Replace with the path to your own image file
IMAGE_PATH = "team_photo.jpg"
image_base64 = encode_image_to_base64(IMAGE_PATH)

# ── STEP 5: Build the API request ────────────────────────────────────────────
# This is like filling out an order form: which image, what do you want to know about it?
analyze_image_details = oci.ai_vision.models.AnalyzeImageDetails(
    # The image itself, encoded in Base64 and sent directly in the request
    image=oci.ai_vision.models.InlineImageDetails(
        source="INLINE",
        data=image_base64
    ),
    # The list of AI tasks you want to run on this image
    features=[
        oci.ai_vision.models.ImageClassificationFeature(
            feature_type="IMAGE_CLASSIFICATION",
            max_results=10       # Return up to 10 labels
        )
    ],
    compartment_id=COMPARTMENT_ID
)

# ── STEP 6: Send the request and get results ─────────────────────────────────
response = ai_vision_client.analyze_image(
    analyze_image_details=analyze_image_details
)

# ── STEP 7: Print the results ────────────────────────────────────────────────
print("🏷️  IMAGE CLASSIFICATION RESULTS:")
print("=" * 45)

for label in response.data.labels:
    # Convert confidence score to percentage
    confidence_pct = label.confidence * 100
    # Create a visual confidence bar (▓ = filled, ░ = empty)
    bar_length    = int(confidence_pct / 5)        # scale to 20 chars
    bar           = "▓" * bar_length + "░" * (20 - bar_length)
    print(f"  {label.name:<25} {bar} {confidence_pct:.1f}%")

Example output (for a photo of a football stadium):

🏷️  IMAGE CLASSIFICATION RESULTS:
=============================================
  Sports                    ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░  97.3%
  Stadium                   ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░  96.8%
  Football                  ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░  92.4%
  Crowd                     ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░  88.9%
  Grass                     ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░  85.2%
  Night                     ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░  79.1%

Clean, ranked labels with a visual confidence bar — easy to read at a glance! 🎯

Part 2 — Object Detection: Find Every Object and Its Location

📌 What the code below does — in plain English:
This code sends the same image but now asks: "Find every object in this photo and tell me exactly where each one is." The result includes the name of each object, its confidence score, and its bounding box coordinates. We also count how many times each type of object appears — useful for inventory counting applications!
import oci
import base64
from collections import Counter

config            = oci.config.from_file()
ai_vision_client  = oci.ai_vision.AIServiceVisionClient(config)
COMPARTMENT_ID    = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."

# Convert image to Base64 (same helper function as before)
def encode_image_to_base64(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

IMAGE_PATH   = "street_scene.jpg"
image_base64 = encode_image_to_base64(IMAGE_PATH)

# ── Build the request — this time asking for OBJECT_DETECTION ────────────────
analyze_image_details = oci.ai_vision.models.AnalyzeImageDetails(
    image=oci.ai_vision.models.InlineImageDetails(
        source="INLINE",
        data=image_base64
    ),
    features=[
        oci.ai_vision.models.ImageObjectDetectionFeature(
            feature_type="OBJECT_DETECTION",
            max_results=50         # Find up to 50 objects in the image
        )
    ],
    compartment_id=COMPARTMENT_ID
)

response = ai_vision_client.analyze_image(
    analyze_image_details=analyze_image_details
)

# ── Display results ───────────────────────────────────────────────────────────
print("📦 OBJECT DETECTION RESULTS:")
print("=" * 65)

object_names = []   # Collect all object names for counting at the end

for obj in response.data.image_objects:
    # Confidence score as a percentage
    conf_pct = obj.confidence * 100

    # Bounding box: normalised coordinates (0.0 to 1.0) relative to image size
    # For a 1000x800 image: left=0.10 means the box starts 100px from the left
    bb = obj.bounding_polygon.normalized_vertices
    print(f"  Object: {obj.name:<20} Confidence: {conf_pct:.1f}%")
    print(f"  Location (normalised): TopLeft({bb[0].x:.2f}, {bb[0].y:.2f}) "
          f"BottomRight({bb[2].x:.2f}, {bb[2].y:.2f})")
    print()

    object_names.append(obj.name)

# ── Print a simple count summary ─────────────────────────────────────────────
print("-" * 65)
print("📊 OBJECT COUNTS:")
object_counts = Counter(object_names)
for name, count in object_counts.most_common():
    print(f"  {name:<25} {count} found")

Example output (for a busy street photo):

📦 OBJECT DETECTION RESULTS:
=================================================================
  Object: Car                  Confidence: 96.4%
  Location (normalised): TopLeft(0.04, 0.51) BottomRight(0.24, 0.89)

  Object: Car                  Confidence: 94.8%
  Location (normalised): TopLeft(0.31, 0.49) BottomRight(0.52, 0.87)

  Object: Person               Confidence: 98.2%
  Location (normalised): TopLeft(0.01, 0.38) BottomRight(0.06, 0.92)

  Object: Person               Confidence: 97.5%
  Location (normalised): TopLeft(0.07, 0.40) BottomRight(0.12, 0.90)

  Object: Traffic Light        Confidence: 91.7%
  Location (normalised): TopLeft(0.43, 0.08) BottomRight(0.47, 0.38)

-----------------------------------------------------------------
📊 OBJECT COUNTS:
  Person                    2 found
  Car                       2 found
  Traffic Light             1 found

Normalised coordinates range from 0.0 to 1.0. This means they work for any image size — you multiply by the actual image width/height to get pixel coordinates. A TopLeft(0.04, 0.51) on a 1000x800 image means pixel (40, 408).

Part 3 — Multi-Feature in One Call: Objects + Text + Faces Together

📌 What the code below does — in plain English:
This is the most powerful pattern: asking OCI Vision to do THREE things in a SINGLE API call — find objects, read any text, AND detect faces — all at once. This saves time (one call instead of three) and costs fewer API credits. Think of it like giving your AI intern a checklist: "Do ALL of these things while you look at the photo!"
import oci
import base64

config            = oci.config.from_file()
ai_vision_client  = oci.ai_vision.AIServiceVisionClient(config)
COMPARTMENT_ID    = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."

def encode_image_to_base64(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

IMAGE_PATH   = "office_scene.jpg"    # A photo of people in an office with signs
image_base64 = encode_image_to_base64(IMAGE_PATH)

# ── Request THREE features in ONE API call ───────────────────────────────────
# This is like giving the AI a checklist of things to look for simultaneously!
analyze_image_details = oci.ai_vision.models.AnalyzeImageDetails(
    image=oci.ai_vision.models.InlineImageDetails(
        source="INLINE",
        data=image_base64
    ),
    features=[
        # Task 1: Find what objects are in the image
        oci.ai_vision.models.ImageObjectDetectionFeature(
            feature_type="OBJECT_DETECTION",
            max_results=20
        ),
        # Task 2: Read any visible text
        oci.ai_vision.models.ImageTextDetectionFeature(
            feature_type="TEXT_DETECTION"
        ),
        # Task 3: Find all human faces
        oci.ai_vision.models.ImageObjectDetectionFeature(
            feature_type="FACE_DETECTION",
            max_results=10
        )
    ],
    compartment_id=COMPARTMENT_ID
)

response = ai_vision_client.analyze_image(
    analyze_image_details=analyze_image_details
)
result   = response.data

# ── Display Object Detection results ─────────────────────────────────────────
print("📦 OBJECTS FOUND:")
print("-" * 40)
if result.image_objects:
    for obj in result.image_objects:
        print(f"  • {obj.name} ({obj.confidence*100:.1f}% confident)")
else:
    print("  No objects detected.")

# ── Display Text Detection results ────────────────────────────────────────────
print("\n🔤 TEXT FOUND:")
print("-" * 40)
if result.image_text:
    for block in result.image_text.blocks:
        print(f"  • \"{block.text}\" ({block.confidence*100:.1f}% confident)")
else:
    print("  No text detected.")

# ── Display Face Detection results ────────────────────────────────────────────
print("\n👤 FACES FOUND:")
print("-" * 40)
if result.detected_faces:
    for i, face in enumerate(result.detected_faces, start=1):
        bb = face.bounding_polygon.normalized_vertices
        print(f"  Face {i}: at position TopLeft({bb[0].x:.2f}, {bb[0].y:.2f})")
    print(f"\n  Total faces: {len(result.detected_faces)}")
else:
    print("  No faces detected.")

Example output (for an office meeting photo with a whiteboard):

📦 OBJECTS FOUND:
----------------------------------------
  • Person (98.2% confident)
  • Person (97.5% confident)
  • Person (94.1% confident)
  • Chair (89.3% confident)
  • Laptop (91.7% confident)
  • Whiteboard (87.4% confident)

🔤 TEXT FOUND:
----------------------------------------
  • "Q3 TARGETS" (97.1% confident)
  • "Revenue: $2.4M" (94.3% confident)
  • "Team Meeting - March 2026" (91.8% confident)

👤 FACES FOUND:
----------------------------------------
  Face 1: at position TopLeft(0.12, 0.15)
  Face 2: at position TopLeft(0.38, 0.13)
  Face 3: at position TopLeft(0.67, 0.16)

  Total faces: 3

Three tasks done in one API call — objects, text, and faces! 🚀

Feature 6: Images from Object Storage (Production Pattern) 📁

In the examples above, we sent images as Base64 directly in the API request. This works great for small images and quick tests. But in production, your images are usually stored in OCI Object Storage buckets — and OCI Vision can read them directly from there!

📌 What the code below does — in plain English:
Instead of reading an image from your local computer and converting it to Base64, this code tells OCI Vision: "The image is already in my Object Storage bucket — go fetch it from there." This is how you build production apps that process thousands of images stored in the cloud.
import oci

config            = oci.config.from_file()
ai_vision_client  = oci.ai_vision.AIServiceVisionClient(config)
COMPARTMENT_ID    = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."

# ── Build the request pointing to an image in Object Storage ─────────────────
# Instead of InlineImageDetails (Base64), we use ObjectStorageImageDetails
analyze_image_details = oci.ai_vision.models.AnalyzeImageDetails(
    image=oci.ai_vision.models.ObjectStorageImageDetails(
        source="OBJECT_STORAGE",                # Tell OCI Vision where to find the image
        namespace_name="your-namespace",        # Your Object Storage namespace
        bucket_name="product-images-bucket",    # The bucket containing the image
        object_name="electronics/laptop_01.jpg" # The specific file path inside the bucket
    ),
    features=[
        oci.ai_vision.models.ImageClassificationFeature(
            feature_type="IMAGE_CLASSIFICATION",
            max_results=5
        ),
        oci.ai_vision.models.ImageObjectDetectionFeature(
            feature_type="OBJECT_DETECTION",
            max_results=20
        )
    ],
    compartment_id=COMPARTMENT_ID
)

response = ai_vision_client.analyze_image(
    analyze_image_details=analyze_image_details
)

print("✅ Analysis complete for: electronics/laptop_01.jpg")
print("\n🏷️  Top Labels:")
for label in response.data.labels:
    print(f"  {label.name}: {label.confidence*100:.1f}%")

print("\n📦 Objects Detected:")
for obj in response.data.image_objects:
    print(f"  {obj.name}: {obj.confidence*100:.1f}%")
✅ When to use Object Storage vs Base64:
Use Base64 when: testing quickly, processing images from a local computer, building a simple prototype
Use Object Storage when: production apps, processing large images (>1MB), batch processing many images, images are already stored in OCI

Feature 7: Batch Image Analysis — Process Thousands at Once 🔁

What Is Batch Analysis?

The analyze_image API we used above works on ONE image at a time. But what if you have 5,000 product photos to analyse? Calling the API 5,000 times in a loop would be slow and inefficient.

OCI Vision's Batch Image Job feature lets you point to an entire Object Storage bucket and say: "Analyse ALL the images in there, please." OCI Vision processes them in parallel behind the scenes and writes all results to an output bucket.

📌 What the code below does — in plain English:
This submits a batch job to OCI Vision: "Please analyse every image in my input bucket and save all the results as JSON files in my output bucket." It's like hiring a team of AI workers to process a huge pile of photos simultaneously. One job = all images processed — no looping required!
import oci
import time

config            = oci.config.from_file()
ai_vision_client  = oci.ai_vision.AIServiceVisionClient(config)
COMPARTMENT_ID    = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
NAMESPACE         = "your-namespace"

# ── Submit a batch image analysis job ────────────────────────────────────────
create_image_job_details = oci.ai_vision.models.CreateImageJobDetails(
    display_name="product-catalogue-analysis-batch",
    compartment_id=COMPARTMENT_ID,

    # Where are the input images stored?
    input_location=oci.ai_vision.models.ObjectStorageLocations(
        location_type="OBJECT_STORAGE_LOCATIONS",
        object_storage_locations=[
            oci.ai_vision.models.ObjectStorageLocation(
                namespace_name=NAMESPACE,
                bucket_name="product-images-input",  # All images in this bucket
                prefix="2026/march/"                 # Only process images in this subfolder
            )
        ]
    ),

    # What should OCI Vision do with each image?
    features=[
        oci.ai_vision.models.ImageClassificationFeature(
            feature_type="IMAGE_CLASSIFICATION",
            max_results=5
        ),
        oci.ai_vision.models.ImageObjectDetectionFeature(
            feature_type="OBJECT_DETECTION",
            max_results=20
        )
    ],

    # Where should the results be saved?
    output_location=oci.ai_vision.models.OutputLocation(
        namespace_name=NAMESPACE,
        bucket_name="analysis-results-output",
        prefix="batch-results/"
    )
)

# Submit the batch job to OCI
job_response = ai_vision_client.create_image_job(
    create_image_job_details=create_image_job_details
)
job_id = job_response.data.id
print(f"✅ Batch job submitted! Job ID: {job_id}")

# ── Wait for the batch job to complete ───────────────────────────────────────
print("⏳ Processing batch (this may take several minutes for large datasets)...")
while True:
    status_response = ai_vision_client.get_image_job(image_job_id=job_id)
    state = status_response.data.lifecycle_state
    print(f"   Status: {state}")
    if state == "SUCCEEDED":
        print(f"\n🎉 Batch complete! Results saved to 'analysis-results-output' bucket.")
        print(f"   Images processed: {status_response.data.image_count}")
        break
    elif state in ("FAILED", "CANCELED"):
        print(f"\n❌ Batch failed. State: {state}")
        break
    time.sleep(30)  # Check every 30 seconds

Example output:

✅ Batch job submitted! Job ID: ocid1.aivisionimgbatch.oc1..aaaaaaaa...
⏳ Processing batch (this may take several minutes for large datasets)...
   Status: ACCEPTED
   Status: IN_PROGRESS
   Status: IN_PROGRESS
   Status: SUCCEEDED

🎉 Batch complete! Results saved to 'analysis-results-output' bucket.
   Images processed: 3,847

Nearly 4,000 images processed automatically — each one producing a JSON result file in the output bucket. This is how large-scale production pipelines work!

Feature 8: Custom Vision Models — Teach OCI Vision YOUR Objects 🏗️

Why Do You Need a Custom Model?

The pre-trained OCI Vision models are trained on millions of common, everyday images. They know what a car, a person, or a tree looks like. But what if you need to detect something very specific to YOUR business?

  • 🍪 A defective biscuit vs a perfect biscuit on your production line?
  • 🩺 A specific type of tumour in an X-ray image?
  • 🌾 A particular crop disease on a wheat field leaf?
  • 🔩 A specific type of manufacturing defect on a circuit board?

The pre-trained model has never seen any of these! You need to train a custom model on YOUR own labelled images. OCI Vision makes this surprisingly easy — no AI expertise required!

HOW CUSTOM MODEL TRAINING WORKS IN OCI VISION:

Step 1: Collect your images
  → 50+ images minimum (more = better accuracy)
  → Example: 100 photos of perfect biscuits + 100 photos of defective biscuits

Step 2: Upload images to OCI Object Storage bucket

Step 3: Use OCI Data Labeling Service to label each image
  → For IMAGE CLASSIFICATION: Tag each image with a class label
    (e.g. "perfect", "cracked", "burnt")
  → For OBJECT DETECTION: Draw bounding boxes on each defect in each image

Step 4: Create a Vision Project in OCI Console
  (a project is just a folder to keep your models organised)

Step 5: Create and Train the model
  → Tell it which dataset to use (your labelled images)
  → Set training duration (OCI Vision recommends 1 hour minimum)
  → Click "Train" — OCI does all the machine learning automatically!

Step 6: Wait for training to complete
  → Image Classification: ~30-60 minutes
  → Object Detection: ~2-4 hours

Step 7: Evaluate the model's accuracy
  → OCI Vision shows you precision, recall, and F1 score
  → Look at the Confusion Matrix to see which classes it gets right/wrong

Step 8: Deploy and use your model!
  → Call the analyze_image API with your custom model's OCID
  → Same code, just one extra parameter: model_id
📌 What the code below does — in plain English:
Once your custom model is trained and deployed, using it is almost identical to using the pre-trained model. The only difference is adding a model_id parameter to your feature request. This tells OCI Vision: "Don't use your default pre-trained model — use MY custom model instead."
import oci
import base64

config            = oci.config.from_file()
ai_vision_client  = oci.ai_vision.AIServiceVisionClient(config)
COMPARTMENT_ID    = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."

# The OCID of YOUR custom trained model
# Find this in OCI Console → Vision → Projects → Your Project → Your Model
CUSTOM_MODEL_OCID = "ocid1.aimodel.oc1..aaaaaaaa...your-custom-model-ocid..."

def encode_image_to_base64(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

# The biscuit image you want to inspect
IMAGE_PATH   = "biscuit_sample_042.jpg"
image_base64 = encode_image_to_base64(IMAGE_PATH)

# ── Use your custom model — only ONE change from the standard code! ───────────
analyze_image_details = oci.ai_vision.models.AnalyzeImageDetails(
    image=oci.ai_vision.models.InlineImageDetails(
        source="INLINE",
        data=image_base64
    ),
    features=[
        oci.ai_vision.models.ImageClassificationFeature(
            feature_type="IMAGE_CLASSIFICATION",
            max_results=3,
            # ↓ THIS is the only change — point to your custom model!
            model_id=CUSTOM_MODEL_OCID
        )
    ],
    compartment_id=COMPARTMENT_ID
)

response = ai_vision_client.analyze_image(
    analyze_image_details=analyze_image_details
)

print(f"🍪 Quality Check Result for: {IMAGE_PATH}")
print("=" * 45)
for label in response.data.labels:
    conf_pct = label.confidence * 100
    icon = "✅" if label.name == "perfect" else "❌"
    print(f"  {icon} {label.name:<20} {conf_pct:.1f}%")

# Automated decision based on the top result
top_result   = response.data.labels[0]
decision     = "PASS — send to packaging" if top_result.name == "perfect" else "FAIL — remove from line"
print(f"\n🏭 Automated Decision: {decision}")

Example output:

🍪 Quality Check Result for: biscuit_sample_042.jpg
=============================================
  ❌ cracked               91.4%
  ❌ burnt                  5.2%
  ✅ perfect                3.4%

🏭 Automated Decision: FAIL — remove from line

The custom model identified this biscuit as cracked with 91.4% confidence. The production system automatically removes it from the belt. No human needed! 🏭

✅ Custom Model Tips for Best Accuracy:
📸 Collect at least 50 images per class (100+ is much better)
⚖️ Keep classes balanced — roughly equal images per class
🌈 Include variety — different lighting, angles, distances in your training images
🔍 For object detection, draw bounding boxes as tight as possible around the object
🔄 After training, review the Confusion Matrix — if one class is confused with another, add more examples of those two classes
❌ Common Custom Model Mistakes:
❌ Too few images — don't expect good accuracy with fewer than 30 images per class
❌ Poor labelling — if your training images are labelled incorrectly, the model learns the wrong thing (garbage in, garbage out!)
❌ No variety — if all your "perfect" images are under bright studio lighting, the model won't work well under dim warehouse lighting
❌ Stopping training too early — give OCI Vision at least 1-2 hours of training time for good results

Real-World Use Cases — Where Is OCI Vision Used ? 

Let's look at actual industries where OCI Vision is transforming work :

  • 🏭 Manufacturing — Quality Control
    Cameras on the production line take photos of every product. OCI Vision custom models check each one for defects (scratches, cracks, wrong colour, missing parts) at speeds no human quality inspector could match — thousands of items per hour.
  • 🛒 Retail — Shelf Management
    Cameras photograph store shelves every hour. OCI Vision Object Detection counts how many of each product is on the shelf. When a product falls below the minimum stock level, a restocking alert is triggered automatically.
  • 🏥 Healthcare — Medical Imaging Assistance
    Custom OCI Vision models trained on medical images (X-rays, MRI scans) help doctors identify potential areas of concern. The AI highlights regions for the doctor to examine more closely. The doctor still makes all decisions — the AI just helps them focus.
  • 🌿 Utilities — Infrastructure Monitoring
    Drones photograph power lines and pylons. OCI Vision detects vegetation growing too close to the lines (fire risk!) or corrosion on the structures — without any human needing to physically inspect miles of lines.
  • 🚗 Smart Parking — Vehicle Counting
    A single camera at the car park entrance photographs each row. OCI Vision counts the vehicles in each photo. A live dashboard shows which sections are full and which have spaces — updated every minute.
  • 🎬 Media — Content Indexing
    A broadcast company has 50 years of video archive. OCI Vision Video Analysis processes all of it and creates a searchable index. Now a researcher can search "find every clip that shows a football match" and get results in seconds.

Quick Summary 📝

What we learned about OCI Vision:

  • OCI Vision → Cloud AI service that analyses images, videos, and documents using deep learning
  • Image Classification → "What is this image about?" Returns labels with confidence scores
  • Object Detection → "What objects are here and WHERE exactly?" Returns bounding box coordinates
  • Face Detection → Finds and locates all human faces in an image with bounding boxes
  • Text Detection (OCR) → Reads any text visible in an image — signs, labels, handwriting
  • Video Analysis → Analyses every frame of a stored video — find what appears and when
  • Multiple Features in One Call → Ask for Classification + Object Detection + Text + Faces simultaneously
  • Base64 vs Object Storage → Base64 for quick tests; Object Storage for production
  • Batch Analysis → Submit one job to process thousands of images in parallel
  • Custom Models → Train OCI Vision on YOUR labelled images to detect YOUR specific objects

You now understand OCI Vision deeply enough to build real production applications with it! 🏆 Computer Vision is one of the fastest-growing areas of AI — and OCI Vision makes it accessible to every developer without needing a PhD. Go build something amazing! 👁️☁️

Comments