Skip to main content

OCI Named Entity Recognition (NER): Extract Entities from Text with Oracle Cloud AI

Calculating read time…

Imagine you are a detective 🕵️ who reads thousands of news articles every day. Your job is to pull out every person's name, every company mentioned, every city visited, every date, and every amount of money — from raw, messy, unstructured paragraphs. Doing that by hand for just 10 articles would take hours. For 10,000 articles? Impossible.

Named Entity Recognition (NER) is the AI superpower that does exactly this — automatically. It reads text and highlights every meaningful "thing" it finds, just like a detective underlining clues in a report. And OCI Language NER lets you use this power through a simple API call — no AI expertise needed!

What Is Named Entity Recognition?

The Simple Explanation

When you read a sentence, your brain automatically recognises that certain words are "special". In the sentence "Sunita Williams flew from Houston to the International Space Station on March 12th", your brain instantly knows:

  • 👤 Sunita Williams = a person's name
  • 📍 Houston = a place
  • 🏢 International Space Station = an organisation / facility
  • 📅 March 12th = a date

You don't have to think hard about this — your brain does it automatically because it learned to recognise these patterns after years of reading. NER teaches a computer to do the exact same thing!

The computer reads a sentence word by word, detects which words are "special entities", and labels each one with a category tag. This process of finding named things and categorising them is called Named Entity Recognition.

💡 Real-world analogy: Think of NER like a super-smart highlighter that reads your documents. As it reads, it highlights every person's name in yellow, every company in blue, every date in green, and every money amount in orange. Then it makes you a neat list of everything it highlighted — sorted by category!

How Does NER Work? 

Under the hood, OCI Language NER uses a process called Natural Language Processing (NLP). Here's how it works step by step — explained as simply as possible:

INPUT TEXT:
"Oracle Corporation, headquartered in Austin Texas, reported $50 billion
 in revenue for fiscal year 2025. CEO Safra Catz announced the results
 at the Oracle CloudWorld conference on 22nd September."

STEP 1: TOKENISATION
Split text into individual words (tokens):
["Oracle", "Corporation", ",", "headquartered", "in", "Austin", "Texas",
 ",", "reported", "$50", "billion", ...]

STEP 2: CONTEXT ANALYSIS
The AI looks at each word AND its surrounding words.
"Oracle Corporation" together → clearly a company name.
"Austin Texas" together → clearly a location.
"$50 billion" → clearly a money amount.
"Safra Catz" → clearly a person's name (especially after "CEO").

STEP 3: ENTITY LABELLING
Each detected entity gets a TYPE label:

"Oracle Corporation"   → ORGANIZATION
"Austin Texas"         → LOCATION (subtype: GPE = Geopolitical Entity)
"$50 billion"          → QUANTITY (subtype: MONEY)
"fiscal year 2025"     → DATE / TIME
"Safra Catz"           → PERSON
"Oracle CloudWorld"    → EVENT / ORGANIZATION
"22nd September"       → DATE / TIME

STEP 4: OUTPUT
Returns structured JSON with each entity, its type, its position in
the text (offset + length), and a confidence score (0 to 1).

The "position" returned is called an offset — it tells you exactly where in the original text each entity starts. Combined with the length, you know the precise span of text for each entity. This is useful for highlighting entities in a document viewer or replacing them in automated pipelines.

All the Entity Types OCI NER Can Detect 🏷️

OCI Language NER can identify 18+ types of entities in text. Let's go through each one with a real example so you understand exactly what each category means:

  • 👤 PERSON — Real people, historical figures, fictional characters
    Example: "Elon Musk visited Narendra Modi in New Delhi"
  • 🏢 ORGANIZATION — Companies, government agencies, teams, institutions
    Example: "Apple and Microsoft both report earnings to the SEC"
  • 📍 LOCATION — Countries, cities, landmarks, rivers, mountains
    Sub-types: GPE (Geopolitical Entity like country/city), FAC (Facility), LOC (non-GPE like a river)
    Example: "The summit was held in Geneva, near Lake Geneva"
  • 📅 DATE / TIME — Calendar dates, durations, times, years
    Example: "The contract starts 1st April 2026 and runs for two years"
  • 💰 QUANTITY — Numbers, percentages, measurements, money
    Sub-types: MONEY, PERCENT, UNIT, CARDINAL (plain numbers), ORDINAL (1st, 2nd)
    Example: "Revenue grew by 23% to reach $4.2 billion"
  • 📧 EMAIL — Email addresses found in text
    Example: "Contact us at support@oracle.com for assistance"
  • 📞 PHONE — Phone numbers in any format
    Example: "Call our helpline at +1-800-555-0199"
  • 🌐 URL — Web addresses found in text
    Example: "Visit www.oracle.com/cloud for more details"
  • 🎭 EVENT — Named events, conferences, historical events
    Example: "The Olympic Games and Oracle CloudWorld both attracted millions"
  • 🎨 WORK_OF_ART — Books, movies, songs, artworks
    Example: "The Great Gatsby was referenced in the keynote speech"
  • ⚖️ LAW — Named laws, regulations, legal acts
    Example: "Compliance with GDPR and SOX regulations is mandatory"
  • 🔬 PRODUCT — Named products, software, hardware
    Example: "Oracle Database 23ai and OCI Generative AI were launched"
💡 What is a subtype?
Some entity types have subtypes for extra precision. For example, LOCATION has subtypes GPE (countries, cities — "Geopolitical Entity"), FAC (facilities like airports and stadiums), and LOC (natural features like mountains and rivers). Subtypes help you be more specific about what kind of entity was found. OCI NER returns both the type AND the subtype in its response!

Why NER Matters — Real Business Impact 🏢

NER isn't just an academic exercise — it powers real products that you probably use every day. Here are concrete examples:

  • 📰 Google News auto-categorisation — When you search for "Apple earnings", Google knows "Apple" here means the company, not the fruit. That's NER in action — understanding context to classify entities correctly.
  • 📞 CRM auto-enrichment — A customer sends an email saying "I met with Sarah from Deloitte in Chicago on Tuesday about the Q2 contract." NER extracts PERSON (Sarah), ORGANIZATION (Deloitte), LOCATION (Chicago), DATE (Tuesday), and automatically fills the CRM fields. No manual data entry!
  • ⚖️ Legal contract review — A law firm processes 500 contracts. NER extracts every party name, every date, every financial term, and every jurisdiction — turning 500 unreadable documents into a searchable database in minutes.
  • 🏥 Medical records processing — A hospital processes patient notes. NER identifies every drug name, every diagnosis, every dosage, and every doctor's name — structuring them for analytics and compliance.
  • 💰 Financial news monitoring — A trading firm monitors 10,000 news articles per hour. NER extracts every company name, every financial figure, and every analyst name — feeding them into automated trading signals.

Setting Up Your Environment — One-Time Setup 🛠️

Step 1: Install the OCI Python SDK

📌 What this does:
This installs Oracle's official Python library onto your computer. Once installed, your Python code can talk directly to all OCI services — including OCI Language NER. You only need to run this once!
pip install oci

Step 2: Create Your OCI API Key

  • Log in to OCI Console at cloud.oracle.com
  • Click your profile avatar (top-right corner) → My Profile
  • Scroll to API Keys → click Add API Key
  • Choose Generate API Key Pair → download the private key file (.pem)
  • Save it as ~/.oci/oci_api_key.pem
  • Copy the configuration snippet shown — you'll use it next

Step 3: Create the OCI Config File

📌 What this file does:
Think of this like your login ID card for OCI. The Python SDK reads this file every time it connects to OCI to verify your identity. Create a file at ~/.oci/config and paste the values you copied from the console.
[DEFAULT]
user=ocid1.user.oc1..aaaaaaaa...your-user-ocid...
fingerprint=aa:bb:cc:dd:ee:ff:00:11:22:33:44:55:66:77:88:99
tenancy=ocid1.tenancy.oc1..aaaaaaaa...your-tenancy-ocid...
region=us-ashburn-1
key_file=~/.oci/oci_api_key.pem

Step 4: Set the Required IAM Policy

📌 What this does:
OCI requires an administrator to explicitly grant permission to use each service. An admin in your account needs to add the line below as a Policy statement in OCI Console → Identity → Policies. Without this, every API call returns "Authorization failed" — even with correct credentials!
allow group <your-group-name> to use ai-service-language-family in tenancy

Your First NER Call — Detect Entities in Any Text! 🐍

The Scenario

You work at a media company. Every hour, hundreds of press releases arrive. You want to automatically extract all the key people, companies, locations, and dates from each one — so they can be indexed, tagged, and searched later. Let's use OCI NER to do exactly that!

Part 1 — Basic NER: Extract All Entities from a Press Release

📌 What the code below does — in plain English:
This code does 4 simple things:
  1. Connects to OCI Language using your credentials
  2. Sends a press release paragraph as text input
  3. Asks OCI NER: "Find every named entity in this text — people, companies, places, dates, money amounts, etc."
  4. Prints each entity found with its type, confidence score, and exact position in the text
Think of it as feeding the paragraph to your AI detective and asking: "What important names and facts are in here?"
import oci

# ── STEP 1: Load your OCI credentials ────────────────────────────────────────
# This reads your ~/.oci/config file and logs you into OCI
config = oci.config.from_file()

# ── STEP 2: Connect to OCI Language service ───────────────────────────────────
# This opens a connection to OCI Language — like dialling the service's number
language_client = oci.ai_language.AIServiceLanguageClient(config)

# ── STEP 3: Write the text you want to analyse ───────────────────────────────
# This could be a press release, a news article, a customer email — any text!
press_release_text = """
Oracle Corporation, headquartered in Austin, Texas, announced record revenue
of $14.3 billion for Q3 fiscal 2026, a 17% increase year-over-year.
CEO Safra Catz presented the results at the Oracle CloudWorld conference
held on 22nd March 2026 in Las Vegas, Nevada.
The company also confirmed a new partnership with Samsung Electronics
to expand OCI services across Southeast Asia by December 2026.
For investor enquiries, contact ir@oracle.com or call +1-650-506-7000.
"""

# ── STEP 4: Build the request ─────────────────────────────────────────────────
# We wrap the text in a TextDocument object — OCI Language needs this format.
# Up to 100 documents can be sent in one API call (batch processing!)
text_document = oci.ai_language.models.TextDocument(
    key="press_release_001",   # A unique key to identify this document
    text=press_release_text,
    language_code="en"          # Language of the text — "en" = English
)

# Build the full request details
batch_detect_entities_details = oci.ai_language.models.BatchDetectLanguageEntitiesDetails(
    documents=[text_document]   # Send one or more documents in a single call
)

# ── STEP 5: Call the OCI NER API ──────────────────────────────────────────────
# This is where the magic happens! Send the text to OCI, get entities back.
response = language_client.batch_detect_language_entities(
    batch_detect_language_entities_details=batch_detect_entities_details
)

# ── STEP 6: Display the results ───────────────────────────────────────────────
print("🔍 NAMED ENTITY RECOGNITION RESULTS")
print("=" * 60)

for document in response.data.documents:
    print(f"\n📄 Document: {document.key}")
    print(f"   Entities found: {len(document.entities)}")
    print("-" * 60)

    # Group entities by type for a cleaner display
    entities_by_type = {}
    for entity in document.entities:
        entity_type = entity.type
        if entity_type not in entities_by_type:
            entities_by_type[entity_type] = []
        entities_by_type[entity_type].append(entity)

    # Print each entity type group
    for entity_type, entities in sorted(entities_by_type.items()):
        print(f"\n  [{entity_type}]")
        for ent in entities:
            sub = f" → {ent.sub_type}" if ent.sub_type else ""
            conf = ent.score * 100
            print(f"    • \"{ent.text}\"{sub}  ({conf:.1f}% confident)")

Output:

🔍 NAMED ENTITY RECOGNITION RESULTS
============================================================

📄 Document: press_release_001
   Entities found: 15
------------------------------------------------------------

  [DATE]
    • "Q3 fiscal 2026"  (98.4% confident)
    • "22nd March 2026"  (99.1% confident)
    • "December 2026"  (97.8% confident)

  [EMAIL]
    • "ir@oracle.com"  (99.9% confident)

  [LOCATION]
    • "Austin, Texas" → GPE  (97.2% confident)
    • "Las Vegas, Nevada" → GPE  (98.6% confident)
    • "Southeast Asia" → LOC  (94.3% confident)

  [ORGANIZATION]
    • "Oracle Corporation"  (99.5% confident)
    • "Oracle CloudWorld"  (91.7% confident)
    • "Samsung Electronics"  (98.1% confident)
    • "OCI"  (96.4% confident)

  [PERSON]
    • "Safra Catz"  (97.9% confident)

  [PHONE]
    • "+1-650-506-7000"  (99.8% confident)

  [QUANTITY]
    • "$14.3 billion" → MONEY  (99.2% confident)
    • "17%" → PERCENT  (98.7% confident)

Brilliant! 🎉 Every important entity in the press release has been found and categorised automatically. Safra Catz is a PERSON, Oracle Corporation is an ORGANIZATION, Las Vegas is a LOCATION, and so on. All done in one API call — in under 1 second!

Part 2 — Understanding Offset and Length: Pinpointing Entities in the Text

Each entity also comes with an offset (where it starts in the text) and a length (how many characters long it is). This is how you can highlight entities in a document viewer or replace them in text processing pipelines.

📌 What the code below does — in plain English:
This code reads the offset and length values from the NER response and uses them to "reconstruct" each entity's location in the original text. It also demonstrates how to replace entity text — useful for anonymisation (replacing names with [PERSON]) or data transformation. Think of offset as "how many characters from the start of the text does this entity begin?"
import oci

config          = oci.config.from_file()
language_client = oci.ai_language.AIServiceLanguageClient(config)

sample_text = ("Priya Sharma from Infosys signed a contract worth $2.5 million "
               "with Amazon Web Services on 15th January 2026 in Bengaluru.")

text_document = oci.ai_language.models.TextDocument(
    key="contract_note",
    text=sample_text,
    language_code="en"
)

response = language_client.batch_detect_language_entities(
    batch_detect_language_entities_details=oci.ai_language.models.BatchDetectLanguageEntitiesDetails(
        documents=[text_document]
    )
)

# ── Show entities with their exact position in the text ──────────────────────
print("📍 ENTITY POSITIONS IN TEXT:")
print(f"\nOriginal text:\n\"{sample_text}\"\n")
print(f"{'Entity':<25 ffset="" ype="">6} {'Length':>7} {'Verified Extraction'}")
print("-" * 80)

for entity in response.data.documents[0].entities:
    offset  = entity.offset
    length  = entity.length
    # Extract the entity text using offset + length from the original string
    # This should match entity.text exactly!
    extracted = sample_text[offset : offset + length]
    print(
        f"{entity.text:<25 entity.type:="" f="" offset:="">6} "
        f"{length:>7} "
        f"\"{extracted}\""
    )

# ── Show how to anonymise text using NER ────────────────────────────────────
print("\n\n🔏 ANONYMISED VERSION (replace entities with type labels):")

anonymised = sample_text
# Process entities in reverse order so offsets stay correct after replacements
entities_sorted = sorted(response.data.documents[0].entities,
                         key=lambda e: e.offset, reverse=True)

for entity in entities_sorted:
    replacement    = f"[{entity.type}]"
    before         = anonymised[:entity.offset]
    after          = anonymised[entity.offset + entity.length:]
    anonymised     = before + replacement + after

print(f"\"{anonymised}\"")

Output:

📍 ENTITY POSITIONS IN TEXT:

Original text:
"Priya Sharma from Infosys signed a contract worth $2.5 million with Amazon Web Services on 15th January 2026 in Bengaluru."

Entity                    Type            Offset  Length  Verified Extraction
--------------------------------------------------------------------------------
Priya Sharma              PERSON               0      12  "Priya Sharma"
Infosys                   ORGANIZATION        18       7  "Infosys"
$2.5 million              QUANTITY            49      12  "$2.5 million"
Amazon Web Services       ORGANIZATION        67      19  "Amazon Web Services"
15th January 2026         DATE                90      17  "15th January 2026"
Bengaluru                 LOCATION           111       9  "Bengaluru"


🔏 ANONYMISED VERSION (replace entities with type labels):
"[PERSON] from [ORGANIZATION] signed a contract worth [QUANTITY] with [ORGANIZATION] on [DATE] in [LOCATION]."

The offset and length values let you extract each entity's exact position in the text. And the anonymisation example shows how you can replace sensitive information with placeholder labels — perfect for privacy compliance or preparing data for analytics! 🔏

Part 3 — Batch NER: Analyse Multiple Documents in One API Call

OCI Language NER supports batch processing — you can send up to 100 text documents in a single API call! This is much more efficient than calling the API 100 times in a loop.

📌 What the code below does — in plain English:
Imagine you receive 5 customer complaint emails and need to extract from each one: who the customer mentioned, what products they named, what dates they referenced, and what locations they mentioned. This code sends all 5 emails in a single API call and gets all entity results back at once. One round trip to the cloud instead of five — much faster!
import oci

config          = oci.config.from_file()
language_client = oci.ai_language.AIServiceLanguageClient(config)

# ── 5 customer complaint emails to process in one batch ──────────────────────
customer_emails = [
    {
        "id":   "email_001",
        "text": "I ordered an iPhone 15 Pro on 3rd March 2026 but it arrived damaged. "
                "Please contact John at Apple Support in Cupertino about order #A98234."
    },
    {
        "id":   "email_002",
        "text": "My account with HDFC Bank in Mumbai was charged $240 twice on 10th Feb. "
                "Please help — my advisor Rahul Verma said he would look into it."
    },
    {
        "id":   "email_003",
        "text": "The AWS training course by Dr. Lisa Chen in Singapore on 28th April 2026 "
                "was cancelled without notice. I paid $1,200 and want a full refund."
    },
    {
        "id":   "email_004",
        "text": "Delivery from Flipkart to my address in Hyderabad was due on 5th January "
                "but never arrived. The tracking shows it left the Delhi warehouse."
    },
    {
        "id":   "email_005",
        "text": "I spoke with Sarah Johnson from Tesla Motors on Tuesday about my Model 3. "
                "She promised a service appointment but I haven't heard back since."
    }
]

# ── Build a list of TextDocument objects — one per email ─────────────────────
text_documents = [
    oci.ai_language.models.TextDocument(
        key=email["id"],
        text=email["text"],
        language_code="en"
    )
    for email in customer_emails
]

# ── Send ALL 5 documents in ONE API call ────────────────────────────────────
response = language_client.batch_detect_language_entities(
    batch_detect_language_entities_details=oci.ai_language.models.BatchDetectLanguageEntitiesDetails(
        documents=text_documents     # All 5 emails sent together
    )
)

# ── Display a summary for each email ─────────────────────────────────────────
print("📬 BATCH NER RESULTS — 5 Customer Emails")
print("=" * 60)

for doc_result in response.data.documents:
    email_index  = int(doc_result.key.split("_")[1]) - 1
    original_text= customer_emails[email_index]["text"]

    print(f"\n📧 {doc_result.key}")
    print(f"   Text preview: \"{original_text[:60]}...\"")

    # Summarise what was found
    persons  = [e.text for e in doc_result.entities if e.type == "PERSON"]
    orgs     = [e.text for e in doc_result.entities if e.type == "ORGANIZATION"]
    locs     = [e.text for e in doc_result.entities if e.type == "LOCATION"]
    dates    = [e.text for e in doc_result.entities if e.type == "DATE"]
    money    = [e.text for e in doc_result.entities if e.type == "QUANTITY" and "MONEY" in (e.sub_type or "")]

    if persons: print(f"   👤 People:        {', '.join(persons)}")
    if orgs:    print(f"   🏢 Organisations: {', '.join(orgs)}")
    if locs:    print(f"   📍 Locations:     {', '.join(locs)}")
    if dates:   print(f"   📅 Dates:         {', '.join(dates)}")
    if money:   print(f"   💰 Money:         {', '.join(money)}")
    print(f"   Total entities: {len(doc_result.entities)}")

Output:

📬 BATCH NER RESULTS — 5 Customer Emails
============================================================

📧 email_001
   Text preview: "I ordered an iPhone 15 Pro on 3rd March 2026 but it arrived dama..."
   👤 People:        John
   🏢 Organisations: Apple Support
   📍 Locations:     Cupertino
   📅 Dates:         3rd March 2026
   Total entities: 5

📧 email_002
   Text preview: "My account with HDFC Bank in Mumbai was charged $240 twice on 10t..."
   👤 People:        Rahul Verma
   🏢 Organisations: HDFC Bank
   📍 Locations:     Mumbai
   📅 Dates:         10th Feb
   💰 Money:         $240
   Total entities: 5

📧 email_003
   Text preview: "The AWS training course by Dr. Lisa Chen in Singapore on 28th Apri..."
   👤 People:        Dr. Lisa Chen
   🏢 Organisations: AWS
   📍 Locations:     Singapore
   📅 Dates:         28th April 2026
   💰 Money:         $1,200
   Total entities: 5

📧 email_004
   Text preview: "Delivery from Flipkart to my address in Hyderabad was due on 5th J..."
   🏢 Organisations: Flipkart
   📍 Locations:     Hyderabad, Delhi
   📅 Dates:         5th January
   Total entities: 4

📧 email_005
   Text preview: "I spoke with Sarah Johnson from Tesla Motors on Tuesday about my Mo..."
   👤 People:        Sarah Johnson
   🏢 Organisations: Tesla Motors
   📅 Dates:         Tuesday
   Total entities: 4

Five emails processed in one API call! 🚀 Each result is perfectly categorised — ready to be stored in a CRM, used for automated routing, or fed into a dashboard showing complaint trends by location, company, and person.

Feature: Filter by Entity Type — Get Only What You Need 🎯

In many real applications, you only care about specific entity types. For example, a financial news system only needs ORGANIZATION and QUANTITY entities. A calendar tool only needs DATE entities. There's no need to process all entity types if you only need a few!

📌 What the code below does — in plain English:
This processes a financial news headline and filters the results to show only ORGANIZATION names and QUANTITY values (money, percentages). Imagine you are building a stock market monitoring tool — you want to know which companies were mentioned and what numbers were associated with them. This code does exactly that filtering in a clean, reusable way.
import oci
from collections import defaultdict

config          = oci.config.from_file()
language_client = oci.ai_language.AIServiceLanguageClient(config)

# A financial news article excerpt
financial_news = """
Microsoft Corporation reported net income of $21.9 billion for Q2 2026,
beating analyst estimates by 12%. The Azure cloud division grew 34% year-on-year,
outperforming Alphabet's Google Cloud which grew 28%. Shares of Microsoft rose 4.7%
to $425 on the Nasdaq, while Amazon Web Services also posted strong results with
$27.4 billion in quarterly revenue, up 19% from last year.
"""

text_doc  = oci.ai_language.models.TextDocument(
    key="fin_news_001",
    text=financial_news,
    language_code="en"
)

response = language_client.batch_detect_language_entities(
    batch_detect_language_entities_details=oci.ai_language.models.BatchDetectLanguageEntitiesDetails(
        documents=[text_doc]
    )
)

entities = response.data.documents[0].entities

# ── Define which entity types we care about ───────────────────────────────────
TYPES_WE_WANT = {"ORGANIZATION", "QUANTITY"}

# ── Filter and group entities by type ─────────────────────────────────────────
filtered = defaultdict(list)
for entity in entities:
    if entity.type in TYPES_WE_WANT:
        filtered[entity.type].append({
            "text":      entity.text,
            "sub_type":  entity.sub_type,
            "confidence": entity.score
        })

# ── Display results in a clean format ─────────────────────────────────────────
print("📈 FINANCIAL NEWS — ENTITY EXTRACTION (Org + Money Only)")
print("=" * 60)

print("\n🏢 COMPANIES / ORGANISATIONS mentioned:")
for item in filtered.get("ORGANIZATION", []):
    print(f"   • {item['text']}  ({item['confidence']*100:.1f}% confident)")

print("\n💰 NUMBERS / FINANCIALS mentioned:")
for item in filtered.get("QUANTITY", []):
    sub = f"[{item['sub_type']}]" if item['sub_type'] else ""
    print(f"   • {item['text']} {sub}  ({item['confidence']*100:.1f}% confident)")

# ── Build a simple company → financials mapping ───────────────────────────────
print("\n\n📊 AUTOMATED INSIGHT: Companies vs Numbers Found Together:")
print("   (Useful for mapping financials to company mentions)\n")

orgs   = [e.text for e in entities if e.type == "ORGANIZATION"]
quants = [e.text for e in entities if e.type == "QUANTITY"]

print(f"   Companies detected: {', '.join(orgs)}")
print(f"   Financial figures:  {', '.join(quants)}")
print(f"\n   Total companies: {len(orgs)} | Total figures: {len(quants)}")

Output:

📈 FINANCIAL NEWS — ENTITY EXTRACTION (Org + Money Only)
============================================================

🏢 COMPANIES / ORGANISATIONS mentioned:
   • Microsoft Corporation  (99.3% confident)
   • Azure  (94.7% confident)
   • Alphabet  (97.2% confident)
   • Google Cloud  (96.8% confident)
   • Microsoft  (98.9% confident)
   • Nasdaq  (95.4% confident)
   • Amazon Web Services  (98.6% confident)

💰 NUMBERS / FINANCIALS mentioned:
   • $21.9 billion [MONEY]  (99.1% confident)
   • 12% [PERCENT]  (98.4% confident)
   • 34% [PERCENT]  (97.8% confident)
   • 28% [PERCENT]  (96.9% confident)
   • 4.7% [PERCENT]  (98.2% confident)
   • $425 [MONEY]  (99.0% confident)
   • $27.4 billion [MONEY]  (99.4% confident)
   • 19% [PERCENT]  (97.5% confident)


📊 AUTOMATED INSIGHT: Companies vs Numbers Found Together:

   Companies detected: Microsoft Corporation, Azure, Alphabet, Google Cloud, Microsoft, Nasdaq, Amazon Web Services
   Financial figures:  $21.9 billion, 12%, 34%, 28%, 4.7%, $425, $27.4 billion, 19%

   Total companies: 7 | Total figures: 8

Feature: NER + PII Detection — Privacy Protection 🔏

NER is also the foundation of PII detection — finding Personally Identifiable Information. PII includes names, phone numbers, email addresses, addresses, account numbers, and other data that could identify a specific individual. Detecting and masking PII is a legal requirement in many industries (GDPR, HIPAA, PCI-DSS).

OCI Language has a dedicated PII Detection and Masking feature built on top of NER. It can automatically find PII entities and replace them with asterisks or placeholder labels.

📌 What the code below does — in plain English:
A customer sends a support message containing their name, credit card number, and phone number. Before this text is stored in a database or sent to an analytics tool, we need to mask all the sensitive personal information. This code calls the OCI PII detection API which finds all PII entities and masks them automatically. The result is a safe version of the text with all sensitive data hidden.
import oci

config          = oci.config.from_file()
language_client = oci.ai_language.AIServiceLanguageClient(config)

# A customer support message containing sensitive PII
support_message = """
Hi Support Team, my name is Aisha Khan and I'm having trouble with my account.
My registered email is aisha.khan@gmail.com and phone number is +44 7911 123456.
My credit card ending in 4782 was charged twice on 10th March 2026.
Please look into this urgently. My account number is AC-789234.
"""

text_doc = oci.ai_language.models.TextDocument(
    key="support_ticket_001",
    text=support_message,
    language_code="en"
)

# ── Configure how to mask the PII ────────────────────────────────────────────
# mode="MASK" replaces PII with asterisks (e.g. "Aisha Khan" → "****** ****")
# mode="REPLACE" replaces PII with a label (e.g. "Aisha Khan" → "[PERSON]")
# mode="REMOVE" deletes the PII entirely

pii_entity_masking = oci.ai_language.models.PiiEntityMask(
    mode="REPLACE",              # Replace PII with type labels
    is_unmasked_from_end=False
)

# Apply this masking rule to all PII entity types
# You can also set different rules per entity type if needed!
masking_config = {
    "ALL": pii_entity_masking    # Apply REPLACE mode to all PII types
}

# ── Call the PII detection API ───────────────────────────────────────────────
pii_response = language_client.batch_detect_language_pii_entities(
    batch_detect_language_pii_entities_details=oci.ai_language.models.BatchDetectLanguagePiiEntitiesDetails(
        documents=[text_doc],
        masking=masking_config
    )
)

# ── Display the results ───────────────────────────────────────────────────────
doc_result = pii_response.data.documents[0]

print("🔏 PII DETECTION AND MASKING RESULTS")
print("=" * 60)

print("\n📋 PII Entities Detected:")
for entity in doc_result.entities:
    print(f"   • Type: {entity.type:<20 60="" code="" doc_result.masked_text="" entity.text="" f="" found:="" labels="" n="" print="" replaced="" safe="" version="" with="">

Output:

🔏 PII DETECTION AND MASKING RESULTS
============================================================

📋 PII Entities Detected:
   • Type: PERSON               Found: "Aisha Khan"
   • Type: EMAIL                Found: "aisha.khan@gmail.com"
   • Type: PHONE_NUMBER         Found: "+44 7911 123456"
   • Type: CREDIT_DEBIT_NUMBER  Found: "4782"
   • Type: DATE_TIME            Found: "10th March 2026"
   • Type: BANK_ACCOUNT_NUMBER  Found: "AC-789234"

✅ SAFE VERSION (PII replaced with labels):
------------------------------------------------------------

Hi Support Team, my name is [PERSON] and I'm having trouble with my account.
My registered email is [EMAIL] and phone number is [PHONE_NUMBER].
My credit card ending in [CREDIT_DEBIT_NUMBER] was charged twice on [DATE_TIME].
Please look into this urgently. My account number is [BANK_ACCOUNT_NUMBER].

All six pieces of personal information were detected and replaced with safe labels! The text is now completely safe to store in analytics systems, share with third-party vendors, or process for business intelligence — with zero risk of data leakage. 🔒

✅ DO: Use OCI PII detection to protect customer data before it enters any analytics pipeline, data warehouse, or third-party system. This is one of the fastest ways to achieve GDPR and HIPAA compliance in your text processing workflows — no manual data governance required!
❌ DON'T: Don't assume that removing someone's name from text makes it fully anonymised. A combination of location + job title + date of birth can still uniquely identify someone even without their name. PII masking is an important first step but may need to be combined with other anonymisation techniques for full compliance.

Feature: Custom NER — Teach OCI to Find YOUR Entities 🏗️

Why Custom NER?

The pre-trained OCI NER model knows about common entity types: people, organisations, locations, dates. But what about entities specific to YOUR industry?

  • 🏥 Healthcare: Drug names like "Metformin 500mg", disease codes like "ICD-10 E11", lab test names
  • ⚖️ Legal: Contract clause types like "Indemnification Clause", case reference numbers
  • 🏭 Manufacturing: Part numbers like "SKU-X7891-B", machine codes, defect categories
  • 💱 Finance: Stock ticker symbols like "ORCL", fund names, regulatory codes

The pre-trained model has never seen these! You need to train a custom NER model on YOUR labelled text. OCI Language makes this possible through its Custom NER feature.

HOW CUSTOM NER TRAINING WORKS IN OCI LANGUAGE:

Step 1: Collect sample text documents
  → At least 50 labelled examples per entity type (more = better)
  → Example: 200 medical notes with drug names manually highlighted

Step 2: Label your data using OCI Data Labeling Service
  → Highlight each entity span in each document
  → Assign the correct entity type label (e.g. "DRUG_NAME", "DOSAGE", "CONDITION")

Step 3: Create a Language Project in OCI Console
  → OCI Console → Analytics & AI → Language → Projects → Create Project
  → A "project" is just a folder to organise your models

Step 4: Create and train the Custom NER model
  → Inside the project: Create Model → choose "Named Entity Recognition"
  → Connect your labelled dataset
  → Set training parameters → click Train
  → OCI Language runs the machine learning automatically!

Step 5: Wait for training to complete (usually 30–120 minutes)

Step 6: Evaluate model accuracy
  → OCI shows Precision, Recall, and F1-Score per entity type
  → Look at the Confusion Matrix to spot weak spots

Step 7: Use your custom model via API
  → Same batch_detect_language_entities() call
  → Just add: endpoint_id=YOUR_CUSTOM_MODEL_ENDPOINT_OCID
✅ Custom NER Tips for Best Results:
📝 Label at least 50 examples per entity type — 100+ is much better
⚖️ Keep entity type counts balanced — don't have 200 DRUG_NAME examples and only 20 DOSAGE examples
🌈 Include variety — different sentence structures, different positions in the text
🔍 Use consistent labelling — if you label "Metformin" as DRUG_NAME, always label it that way
📊 Review the F1-score per entity type — if one type scores below 0.70, add more training examples for it
❌ Common Custom NER Mistakes:
❌ Inconsistent labelling — sometimes labelling "Dr. Smith" as PERSON, sometimes ignoring it
❌ Too few examples — don't expect a custom model to work well with fewer than 30 examples
❌ Overlapping entity types — if DRUG_NAME and MEDICATION_CODE overlap in meaning, merge them into one type
❌ Not evaluating before deploying — always check the Precision and Recall scores before using the model in production

Real-World Use Cases — NER in Production 

Here are real examples of how NER is transforming industries:

  • 📰 Media Intelligence — Automatic Article Tagging
    A news agency publishes 2,000 articles per day. OCI NER processes every article as it's published, extracts all PERSON, ORGANIZATION, and LOCATION entities, and automatically adds them as searchable tags. Readers can search "all articles mentioning Tesla" and get instant results — no human tagger required.
  • 📞 Customer Support Automation
    A telecom company processes 50,000 support tickets per day. OCI NER extracts the product name, issue date, and location from each ticket. Tickets are automatically routed to the correct team based on the product entity found. Response time drops from 4 hours to 12 minutes.
  • ⚖️ Contract Intelligence at Law Firms
    A law firm uploads 10,000 supplier contracts. Custom OCI NER trained on legal text extracts every party name, jurisdiction, payment term, and expiry date. A searchable contract database is built automatically — lawyers can instantly find all contracts with a specific supplier or contracts expiring in Q2 2027.
  • 🏥 Medical Record Structuring
    A hospital has 20 years of free-text doctor notes. Custom OCI NER trained on clinical text extracts diagnoses, medications, dosages, and procedures. These structured entities feed a clinical analytics platform — identifying treatment patterns and outcomes.
  • 💰 Investment Research Acceleration
    An investment fund monitors 5,000 earnings calls per quarter. OCI NER extracts every company name, financial figure, and forward-looking date from every transcript. Analysts have a searchable database of all numbers ever mentioned in earnings calls — in seconds instead of weeks.

Quick Summary 📝

What we learned about OCI Named Entity Recognition:

  • NER → Reads text and automatically finds + categorises all named "things" — people, organisations, locations, dates, money amounts, emails, phones, and more
  • Entity Types → OCI NER detects 18+ types including PERSON, ORGANIZATION, LOCATION, DATE, QUANTITY, EMAIL, PHONE, URL, EVENT, LAW, and PRODUCT — each with optional subtypes for extra precision
  • Offset + Length → Every entity comes with its exact position in the text — useful for highlighting, replacing, and processing specific spans
  • Confidence Score → Every entity gets a 0–1 confidence score. Use a threshold (e.g., 0.80) to filter uncertain detections
  • Batch Processing → Send up to 100 documents in a single API call — much faster than processing one at a time
  • PII Detection + Masking → Built-on-NER feature that finds sensitive personal info and can mask, replace, or remove it automatically — key for GDPR compliance
  • Custom NER → Train OCI Language on YOUR labelled text to detect YOUR industry-specific entity types (drug names, part numbers, legal clauses, stock tickers, etc.)

You now understand OCI NER deeply enough to build real production applications with it! 🏆 NER is one of the most practically useful AI capabilities — it turns unstructured text into structured, searchable, actionable data. Every company sitting on thousands of emails, documents, and reports has a gold mine waiting to be extracted. OCI NER is the pickaxe! Happy building! 🔍☁️

Comments