Imagine a doctor who records a 45-minute patient consultation on their phone every day. At the end of the week that's over 4 hours of audio they need to type up manually. What if a computer could listen to all of it and write every single word down automatically — in under 5 minutes?
That's exactly what OCI Speech does! 🎤 It listens to your audio or video files and converts every spoken word into accurate, timestamped, readable text — with zero AI expertise required from you.
What Is OCI Speech?
Think of OCI Speech like a super-smart robot secretary. You give it an audio or video file. It listens very carefully. Then it writes down everything it heard — perfectly.
Under the hood, OCI Speech uses a technology called ASR — Automatic Speech Recognition. ASR is the same kind of AI that powers Siri, Google Assistant, and Alexa. But OCI Speech is built specifically for businesses — it handles noisy recordings, different accents, multiple speakers, and dozens of languages.
💡 Real-world analogy: Think of ASR like this — you have a really talented friend who has listened to millions of hours of spoken English their whole life. When you play them any audio clip, they can write it down word for word without missing a beat. OCI Speech is that friend, but available 24/7 and able to process hours of audio in minutes!
What Can OCI Speech Actually Do? 🚀
OCI Speech is not just a simple recording-to-text tool. It has several powerful features:
- 🎤 Batch Speech-to-Text (File Transcription) — Upload an audio or video file, get back a text transcript. Perfect for recorded meetings, podcasts, lectures, customer service calls.
- ⚡ Real-Time Speech Transcription — Live, as-you-speak transcription via WebSocket connection. Words appear on screen as you talk. Perfect for live captioning, voice assistants, voice-controlled apps.
- 🌍 Multi-Language Support — Oracle's own ASR model supports English, Spanish, and Portuguese. The Whisper model (also available!) supports 50+ languages including Hindi, French, German, Japanese, and more.
- 🧑🤝🧑 Speaker Diarization — Identifies WHO said WHAT. In a meeting with 5 people, it labels each sentence with the correct speaker. Like putting name tags on every sentence!
- ⏱️ Timestamps on Every Word — Every single word gets a start time and end time. So you can search for any word and jump to exactly that moment in the audio.
- 🔤 Text Normalization — Automatically converts spoken numbers, dates, and addresses into proper written form. "three hundred and forty two dollars" becomes "$342" automatically.
- 🚫 Profanity Filtering — Can detect bad words and either mask them (replace with ***), remove them entirely, or just tag them for review.
- 📊 Confidence Scores — Each word gets a score between 0 and 1 showing how confident the AI is that it heard that word correctly. A score of 0.98 means almost certain. A score of 0.45 means uncertain.
- 📝 SRT File Output — Generates subtitle files (the same .srt format used by YouTube and Netflix) so you can add captions to your videos automatically!
Let's go deeper into the most important features one by one. 🎯
Feature 1: Batch Transcription — Audio Files to Text 📁
How It Works (Step by Step)
Batch transcription is like dropping a box of cassette tapes on your robot secretary's desk and saying "type all of these up, please!" The robot takes them one by one, listens carefully, and produces a neat text file for each one.
Here's the exact flow inside OCI:
YOUR AUDIO FILE (mp3 / wav / mp4 / m4a / ogg / flac / aac)
│
│ Step 1: Upload to OCI Object Storage (your cloud file cabinet)
▼
┌─────────────────────────────────┐
│ OCI Object Storage Bucket │
│ e.g. "my-audio-bucket" │
│ 📁 meeting_recording.mp3 │
└─────────────────┬───────────────┘
│
│ Step 2: Create a Transcription Job
│ (tell OCI Speech: "please process this file")
▼
┌─────────────────────────────────────────────────────────────┐
│ OCI Speech Service │
│ │
│ 🧠 Pre-trained ASR Model (Oracle) OR Whisper (OpenAI) │
│ │
│ Processes audio → produces text │
│ Adds timestamps to every word │
│ Applies normalization, profanity filter, confidence scores │
└─────────────────────────────────┬───────────────────────────┘
│
│ Step 3: Output written back to Object Storage
▼
┌─────────────────────────────────────────────────────────────┐
│ OCI Object Storage Bucket (Output) │
│ 📄 meeting_recording.json ← Full transcript with data │
│ 📄 meeting_recording.srt ← Subtitle file (optional) │
└─────────────────────────────────────────────────────────────┘
│
│ Step 4: Your app reads the results!
▼
YOUR APPLICATION 🎉
The most important thing to understand: OCI Speech reads files from — and writes results to — OCI Object Storage. Object Storage is just Oracle's cloud file cabinet. Before you can use OCI Speech, your audio file must be sitting in a bucket there.
A "Job" is just a batch request you send to OCI Speech. You say: "Here is my audio file in this Object Storage bucket. Please transcribe it and save the result to this other bucket." One job can contain up to 100 audio files processed at the same time. You can submit a job via the OCI Console (point-and-click), the OCI CLI (command line), the REST API, or the Python SDK.
Supported Audio and Video Formats
- 🎵 Audio: MP3, WAV, OGG, FLAC, AAC, M4A, AC3, AMR, ALAW, MULAW
- 🎬 Video: MP4, MKV, MOV, AVI, WebM
- 📏 Max file size: 1 GB per file
- ⏳ Max duration: Up to 4 hours per file
Trying OCI Speech in the Console — No Code Needed! 🖥️
Before we write any code, let's try OCI Speech directly from the browser. This is the easiest way to see what it does!
- Step 1: Log in to OCI Console at
cloud.oracle.com - Step 2: Click the hamburger menu (top-left ☰) → Analytics & AI → Speech
- Step 3: Click Create Transcription Job
-
Step 4: Fill in the job details:
- Job Name: Give it any name, e.g.
my-first-speech-job - Compartment: Choose your compartment (your account's folder)
- Language: Choose the language spoken in the audio (e.g. English US)
- Model: Choose Oracle ASR (for English/Spanish/Portuguese) or Whisper (for 50+ languages)
- Output Format: Choose JSON (always included) and optionally SRT for subtitles
- Input Location: Select your Object Storage bucket containing the audio file
- Output Location: Select or create an Object Storage bucket for results
- Job Name: Give it any name, e.g.
- Step 5: Click Create — the job starts immediately!
- Step 6: Wait for the status to change to SUCCEEDED (a 10-minute recording usually finishes in under 2 minutes!)
- Step 7: Download the output JSON or SRT file from your output bucket 🎉
Understanding the Output — What Does the JSON Look Like? 📄
After the job completes, OCI Speech writes a JSON file to your output bucket. JSON is just a structured text format that computers (and humans) can read easily. Let's look at what the output actually contains:
{
"transcriptions": [
{
"transcription": "Hello everyone welcome to today's meeting about the new product launch.",
"confidence": 0.97,
"tokens": [
{
"token": "Hello",
"confidence": "0.99",
"startTime": "PT0.50S",
"endTime": "PT0.90S"
},
{
"token": "everyone",
"confidence": "0.98",
"startTime": "PT0.92S",
"endTime": "PT1.40S"
},
{
"token": "welcome",
"confidence": "0.72",
"startTime": "PT1.42S",
"endTime": "PT1.90S"
}
...
]
}
]
}
Let's understand what each part means:
- 📝 "transcription" — The full sentence written out as readable text
- 📊 "confidence" (overall) — How confident the AI is about the whole sentence. 0.97 = 97% sure. Excellent!
- 🔤 "token" — Each individual word
- 📊 "confidence" (per word) — How sure the AI is about that specific word. "welcome" at 0.72 means it was slightly uncertain — maybe there was background noise at that moment
- ⏱️ "startTime" / "endTime" — When that word was spoken. "PT1.42S" means "at 1.42 seconds into the audio". This lets you jump to any word in the audio instantly!
💡 Think of confidence like a school grade: 0.95-1.00 = A (very confident), 0.80-0.94 = B (fairly confident), 0.70-0.79 = C (uncertain), below 0.70 = D (might be wrong — review manually!)
Feature 2: Speaker Diarization — Who Said What? 🧑🤝🧑
Imagine you record a board meeting with 5 executives all discussing the quarterly results. Without diarization, the transcript is just one long blob of text — you can't tell who said what. With Speaker Diarization, OCI Speech labels each sentence with a speaker number!
WITHOUT Diarization (confusing!): ────────────────────────────────────────────────────────────── "Good morning everyone. Let's start with the Q3 numbers. Revenue was up 12%. That's amazing! What drove that growth? Mainly the new product line and better marketing strategy." WITH Diarization (clear! 🎉): ────────────────────────────────────────────────────────────── Speaker 1: "Good morning everyone. Let's start with the Q3 numbers." Speaker 2: "Revenue was up 12%." Speaker 3: "That's amazing! What drove that growth?" Speaker 1: "Mainly the new product line and better marketing strategy."
OCI Speech can identify between 2 and 16 speakers in a single recording.
- You can either tell it how many speakers there are (e.g. "this meeting had 4 people")
- Or you can let OCI Speech automatically detect the number of speakers — no prior knowledge needed!
🏥 Healthcare: Doctor-patient consultation → label "Doctor:" and "Patient:" separately for medical records
📞 Call Centers: Agent vs Customer → analyse agent performance, customer sentiment separately
⚖️ Legal: Court transcriptions → label each witness, lawyer, judge clearly
🎙️ Podcasts/Interviews: Automatically separate host from guest(s)
Feature 3: Whisper Model — 50+ Languages Support 🌍
Oracle's own ASR model is excellent for English, Spanish, and Portuguese. But what if your audio is in Hindi, French, German, Japanese, or Korean?
OCI Speech now includes the Whisper model — an open-source model originally created by OpenAI, now hosted and managed securely inside OCI. Whisper was trained on a massive amount of multilingual audio from across the internet and supports over 50 languages.
- 🌐 Supported languages include: English, Spanish, Portuguese, French, German, Italian, Dutch, Russian, Japanese, Korean, Chinese (Mandarin), Arabic, Hindi, Turkish, and 36+ more
- 🔢 Whisper has 5 sizes:
tiny,base,small,medium,large-v2. OCI provides the medium size as the best balance of speed vs accuracy - 🧑🤝🧑 Whisper on OCI also supports speaker diarization
- 🔌 Uses the exact same API, SDK, and Console interface as the Oracle ASR model — just select "Whisper" in the model dropdown!
💡 Oracle ASR vs Whisper — which to choose?
- Use Oracle ASR → if your audio is English, Spanish, or Portuguese and you need the highest accuracy
- Use Whisper → if your audio is in any other language, or you need multilingual support
Feature 4: Real-Time Speech Transcription ⚡
Batch transcription processes a completed audio file after the fact. But what if you need the text to appear as someone is speaking right now? That's what Real-Time Speech Transcription is for!
Think of it like subtitles on a live TV broadcast — words appear at the bottom of the screen within milliseconds of the presenter speaking them. OCI's real-time service works the same way using a WebSocket connection.
HOW REAL-TIME TRANSCRIPTION WORKS:
Your Microphone 🎙️
│
│ streams audio chunks (every few milliseconds)
▼
┌──────────────────────────────────────────────────────┐
│ OCI Real-Time Speech Service │
│ │
│ Receives audio chunks → processes continuously │
│ Returns partial results immediately ("Hello eve-") │
│ Returns final confirmed results ("Hello everyone") │
└──────────────────────────────┬───────────────────────┘
│
│ sends back text results instantly
▼
Your Application 💻
(displays words as they arrive)
Real-time uses a WebSocket connection — a persistent two-way link between your app and OCI. Think of it like a telephone call: the line stays open and both sides can talk continuously, as opposed to sending letters back and forth (which is what the batch API does).
- 💬 Partial results: As you're still speaking, interim text appears (may change as the AI gets more context)
- ✅ Final results: Once a sentence is complete, the final confirmed text is locked in
- 🌍 Languages: English, Spanish, Portuguese (Oracle ASR). Whisper multilingual support also available in real-time!
🎓 Live lecture captions — Students who are hard of hearing can follow along instantly
📞 Live customer service — Agent's conversation is transcribed live for supervisor monitoring
🤖 Voice assistants — "Hey app, book me a meeting for tomorrow" → app acts on spoken command
🌐 Live translation — Combine with OCI Language translation for real-time speech-to-foreign-text!
Setting Up Your Environment — Before Any Code 🛠️
Before we write Python code to use OCI Speech, we need to do a few one-time setup steps. Think of this like setting up your toolbox before starting a building project!
Step 1: Install the OCI Python SDK
This installs the official Oracle Python library onto your computer. It's like downloading an app — once installed, you can use all of OCI's services from Python. You only need to do this once!
pip install oci
Step 2: Create an API Key in OCI Console
- Log in to OCI Console → click your profile icon (top-right) → My Profile
- Scroll down to API Keys → click Add API Key
- Select Generate API Key Pair → Download both the Private Key (.pem file) and copy the config snippet shown
- Save the private key file somewhere safe on your computer (e.g.
~/.oci/oci_api_key.pem)
Step 3: Create the OCI Config File
The OCI config file is like your ID card for OCI. It tells the Python SDK who you are, which account you belong to, and where your secret private key is stored. Without this, OCI won't know who is making the API call!
Create a file at ~/.oci/config with this content (paste your own values from the console):
[DEFAULT] user=ocid1.user.oc1..aaaaaaaa...your-user-ocid-here... fingerprint=aa:bb:cc:dd:ee:ff:00:11:22:33:44:55:66:77:88:99 tenancy=ocid1.tenancy.oc1..aaaaaaaa...your-tenancy-ocid-here... region=us-ashburn-1 key_file=~/.oci/oci_api_key.pem
Each line in this file tells OCI something important:
- user — Your unique user ID in OCI (starts with
ocid1.user...) - fingerprint — A short code that identifies your specific API key
- tenancy — Your company's OCI account ID (starts with
ocid1.tenancy...) - region — Which OCI data center to use (e.g.
us-ashburn-1for US East,ap-mumbai-1for India) - key_file — Where your private key file lives on your computer
Step 4: Set Up IAM Policy (Admin Task)
OCI requires you to explicitly grant permission to use each service. An administrator in your account needs to add this policy. Think of it like a building security guard — even with your ID, you still need the right clearance level!
This grants your user group permission to use OCI Speech and access Object Storage. Without these policies, every API call will return "Authorization Failed". Ask your OCI account administrator to run these commands — or if you are the admin, do it yourself!
allow group <your-group-name> to manage ai-service-speech-family in tenancy allow group <your-group-name> to manage object-family in tenancy allow group <your-group-name> to read tag-namespaces in tenancy
Your First Python Code — Create a Transcription Job! 🐍
The Scenario
Let's say you have recorded a team meeting and saved the audio file as team_meeting.mp3.
You've uploaded it to an OCI Object Storage bucket called audio-input-bucket.
You want OCI Speech to transcribe it and save the result to a bucket called transcript-output-bucket.
Part 1 — Connect to OCI Speech and Create the Job
This code does 3 things:
- Reads your OCI config file to log in to your OCI account
- Creates a connection to the OCI Speech service
- Sends a request: "Please transcribe the file
team_meeting.mp3from bucketaudio-input-bucketand save the result totranscript-output-bucket"
import oci
import time
# ── STEP 1: Load your OCI credentials from the config file ──────────────────
# This reads your ~/.oci/config file and logs you in to OCI.
config = oci.config.from_file()
# ── STEP 2: Create a Speech client ──────────────────────────────────────────
# This opens a connection to the OCI Speech service.
# Think of it like dialling the Speech service's phone number.
speech_client = oci.ai_speech.AIServiceSpeechClient(config)
# ── STEP 3: Set your compartment ID ─────────────────────────────────────────
# The compartment is like a folder in OCI. All your resources live inside one.
# Find your compartment OCID in OCI Console → Identity → Compartments
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
# ── STEP 4: Define where your audio file is ─────────────────────────────────
# Tell OCI which bucket and filename to read the audio from.
INPUT_NAMESPACE = "your-object-storage-namespace" # found in Console → Object Storage
INPUT_BUCKET_NAME = "audio-input-bucket"
INPUT_FILE_NAME = "team_meeting.mp3"
# ── STEP 5: Define where to save the transcript ─────────────────────────────
OUTPUT_BUCKET_NAME = "transcript-output-bucket"
# ── STEP 6: Build the transcription job request ─────────────────────────────
# This is like filling out an order form for OCI Speech.
transcription_job_details = oci.ai_speech.models.CreateTranscriptionJobDetails(
# Give your job a friendly name so you can find it in the console
display_name="my-team-meeting-transcript",
# Which OCI account folder (compartment) to put this job in
compartment_id=COMPARTMENT_ID,
# Describe the input — where is the audio file?
input_location=oci.ai_speech.models.ObjectListInlineInputLocation(
location_type="OBJECT_LIST_INLINE_INPUT_LOCATION",
object_locations=[
oci.ai_speech.models.ObjectLocation(
namespace_name=INPUT_NAMESPACE,
bucket_name=INPUT_BUCKET_NAME,
object_names=[INPUT_FILE_NAME]
)
]
),
# Describe the output — where should the transcript be saved?
output_location=oci.ai_speech.models.OutputLocation(
namespace_name=INPUT_NAMESPACE,
bucket_name=OUTPUT_BUCKET_NAME,
prefix="transcripts/" # Saves inside a "transcripts/" folder in the bucket
),
# Configure the AI model settings
model_details=oci.ai_speech.models.TranscriptionModelDetails(
domain="GENERIC", # "GENERIC" works for most use cases
language_code="en-US" # Language of the audio: English (US)
),
# Optional: also generate an SRT subtitle file
additional_transcription_formats=["SRT"],
# Optional: enable profanity filtering (MASK replaces bad words with ***)
normalization=oci.ai_speech.models.TranscriptionNormalization(
is_punctuation_enabled=True, # Add automatic punctuation
filters=[
oci.ai_speech.models.ProfanityTranscriptionFilter(
type="PROFANITY",
mode="MASK" # Options: MASK, REMOVE, TAG
)
]
)
)
# ── STEP 7: Submit the job ───────────────────────────────────────────────────
# This sends the request to OCI. The job starts processing in the cloud!
response = speech_client.create_transcription_job(
create_transcription_job_details=transcription_job_details
)
# Save the Job ID so we can check its status later
job_id = response.data.id
print(f"✅ Job submitted successfully! Job ID: {job_id}")
print(f" Status: {response.data.lifecycle_state}")
Output after running:
✅ Job submitted successfully! Job ID: ocid1.aispeechtranscriptionjob.oc1..aaaaaaaa... Status: ACCEPTED
The status ACCEPTED means OCI received your request and added it to the processing queue.
It will move through: ACCEPTED → IN_PROGRESS → SUCCEEDED 🎉
Part 2 — Wait for the Job to Complete
After submitting the job, we need to wait for OCI Speech to finish processing. This code checks the job status every 15 seconds and prints an update. As soon as the status changes to "SUCCEEDED" or "FAILED", it stops waiting and tells us the result. Think of it like repeatedly checking your phone to see if the pizza delivery app says "Order Delivered!"
# ── Poll the job status until it completes ───────────────────────────────────
print("\n⏳ Waiting for transcription to complete...")
while True:
# Ask OCI: "What's the current status of this job?"
job_status_response = speech_client.get_transcription_job(
transcription_job_id=job_id
)
current_state = job_status_response.data.lifecycle_state
print(f" Current status: {current_state}")
# If the job is done (success or failure), stop waiting
if current_state == "SUCCEEDED":
print("\n🎉 Transcription complete! Results saved to your output bucket.")
break
elif current_state in ("FAILED", "CANCELED"):
print(f"\n❌ Job did not complete. Final status: {current_state}")
print(f" Details: {job_status_response.data.lifecycle_details}")
break
# Wait 15 seconds before checking again
time.sleep(15)
Output while waiting:
⏳ Waiting for transcription to complete... Current status: ACCEPTED Current status: IN_PROGRESS Current status: IN_PROGRESS Current status: SUCCEEDED 🎉 Transcription complete! Results saved to your output bucket.
Part 3 — Read and Display the Transcript
Now that the transcript JSON file is sitting in our output bucket, this code downloads it from Object Storage and reads the text. It then prints the full transcript and flags any words where the AI was less than 80% confident — those are words you might want to double-check manually!
import json
# ── Download the transcript JSON from Object Storage ─────────────────────────
# First, create a connection to Object Storage
object_storage_client = oci.object_storage.ObjectStorageClient(config)
# The output file follows a predictable naming pattern:
# prefix/namespace_bucket_filename.json
output_file_key = f"transcripts/{INPUT_NAMESPACE}_{INPUT_BUCKET_NAME}_{INPUT_FILE_NAME}.json"
# Download the file from the output bucket
get_object_response = object_storage_client.get_object(
namespace_name=INPUT_NAMESPACE,
bucket_name=OUTPUT_BUCKET_NAME,
object_name=output_file_key
)
# Read the file content and parse it as JSON
transcript_data = json.loads(get_object_response.data.text)
# ── Print the full transcript ─────────────────────────────────────────────────
full_transcript = transcript_data["transcriptions"][0]["transcription"]
overall_confidence = transcript_data["transcriptions"][0]["confidence"]
print("=" * 60)
print("📄 FULL TRANSCRIPT:")
print("=" * 60)
print(full_transcript)
print(f"\n📊 Overall Confidence Score: {overall_confidence:.2f} ({overall_confidence*100:.1f}%)")
# ── Flag low-confidence words (below 80%) for manual review ──────────────────
print("\n⚠️ Low-confidence words to review manually:")
print("-" * 50)
low_confidence_found = False
for word in transcript_data["transcriptions"][0]["tokens"]:
# Skip punctuation marks (they don't have confidence scores)
if word["token"] in (".", ",", "!", "?", ":", ";"):
continue
# Check if confidence is below 80%
if float(word["confidence"]) < 0.80:
low_confidence_found = True
print(
f" Word: '{word['token']:<15}' "
f"| Confidence: {float(word['confidence']):.2f} "
f"| Timestamp: {word['startTime']}"
)
if not low_confidence_found:
print(" None! The AI was highly confident about every word. 🌟")
Example output:
============================================================ 📄 FULL TRANSCRIPT: ============================================================ Good morning everyone. Let's start with the Q3 results. Revenue was up 12% compared to last year. The new product line contributed approximately $342,000 in additional sales. We need to discuss the marketing strategy going forward. 📊 Overall Confidence Score: 0.94 (94.0%) ⚠️ Low-confidence words to review manually: -------------------------------------------------- Word: 'approximately ' | Confidence: 0.76 | Timestamp: PT18.30S Word: 'strategy ' | Confidence: 0.71 | Timestamp: PT34.10S
94% overall confidence is excellent! 🌟 Two words were flagged as uncertain — you can jump to those timestamps in the audio to verify.
Using the Whisper Model — 50+ Languages 🌍
Want to transcribe a meeting recorded in Hindi or French?
Just change the model_details section of your job request!
This shows how to swap the Oracle ASR model for the Whisper model. The only change is in the
model_details section — everything else stays exactly the same!
Here we also enable Speaker Diarization to label who said what.
# Use Whisper model for multilingual transcription with speaker diarization
model_details_whisper = oci.ai_speech.models.TranscriptionModelDetails(
# Tell OCI to use the Whisper model instead of the Oracle ASR model
model_type="WHISPER_MEDIUM",
# Language of the audio: set to "auto" to let Whisper auto-detect!
# Or specify: "fr" for French, "de" for German, "hi" for Hindi, "ja" for Japanese...
language_code="auto",
transcription_settings=oci.ai_speech.models.WhisperModelTranscriptionSettings(
# Enable speaker diarization: identify who said what
diarization=oci.ai_speech.models.Diarization(
is_diarization_enabled=True, # Turn on the feature
# Optional: tell Whisper exactly how many speakers to look for.
# If you leave this out, Whisper auto-detects the number!
number_of_speakers=3 # e.g. 3 people in the meeting
)
)
)
# Now create the job exactly as before, just using model_details_whisper
transcription_job_details_whisper = oci.ai_speech.models.CreateTranscriptionJobDetails(
display_name="multilingual-meeting-with-diarization",
compartment_id=COMPARTMENT_ID,
input_location=oci.ai_speech.models.ObjectListInlineInputLocation(
location_type="OBJECT_LIST_INLINE_INPUT_LOCATION",
object_locations=[
oci.ai_speech.models.ObjectLocation(
namespace_name=INPUT_NAMESPACE,
bucket_name=INPUT_BUCKET_NAME,
object_names=["french_team_call.mp3"] # Your multilingual file
)
]
),
output_location=oci.ai_speech.models.OutputLocation(
namespace_name=INPUT_NAMESPACE,
bucket_name=OUTPUT_BUCKET_NAME,
prefix="transcripts/"
),
model_details=model_details_whisper # <-- This is the only key change!
)
response = speech_client.create_transcription_job(
create_transcription_job_details=transcription_job_details_whisper
)
print(f"✅ Whisper multilingual job submitted! ID: {response.data.id}")
That's it! The same code, just a different model_type.
Whisper will automatically detect the language (if you set language_code="auto"),
transcribe it, and label each speaker. 🎉
Reading Diarization Results — Who Said What? 🧑🤝🧑
After a diarization-enabled job finishes, the JSON output contains a "speakerLabel" field on every word. This code reads the transcript and prints a nicely formatted conversation view — showing exactly which speaker said which sentence. Like reading a movie script with character names!
import json
# ── Read the diarized transcript JSON ─────────────────────────────────────────
# (Assume we've already downloaded it from Object Storage as shown earlier)
with open("diarized_transcript.json", "r") as f:
diarized_data = json.load(f)
# ── Display the conversation with speaker labels ──────────────────────────────
print("=" * 60)
print("🎙️ CONVERSATION TRANSCRIPT WITH SPEAKER LABELS:")
print("=" * 60)
current_speaker = None # Track who is speaking right now
current_sentence = [] # Collect words into sentences
for token in diarized_data["transcriptions"][0]["tokens"]:
word = token.get("token", "")
speaker = token.get("speakerLabel", "Speaker ?")
timestamp = token.get("startTime", "")
# Detect when the speaker changes
if speaker != current_speaker:
# Print the previous speaker's sentence before switching
if current_sentence and current_speaker:
print(f"\n [{current_speaker}]: {' '.join(current_sentence)}")
# Start a new sentence for the new speaker
current_speaker = speaker
current_sentence = [word]
else:
# Same speaker — keep building the sentence
current_sentence.append(word)
# Don't forget to print the last sentence!
if current_sentence and current_speaker:
print(f"\n [{current_speaker}]: {' '.join(current_sentence)}")
print("\n" + "=" * 60)
Example output:
============================================================ 🎙️ CONVERSATION TRANSCRIPT WITH SPEAKER LABELS: ============================================================ [Speaker 1]: Good morning everyone let's begin the meeting . [Speaker 2]: Thanks for joining today's session on the new product launch strategy . [Speaker 1]: Before we start I want to confirm everyone received the briefing document . [Speaker 3]: Yes I reviewed it last night and have several questions about the timeline . [Speaker 2]: Great we'll cover the timeline in detail in the second half . ============================================================
Beautiful! 🎉 You can see exactly who said what, in order. Now imagine doing this for a 2-hour board meeting with 8 executives — OCI Speech handles it all automatically!
Checking Word-Level Confidence Scores 📊
One of OCI Speech's most powerful features is giving you a confidence score for EVERY individual word. Here's a quick practical example of how to use this for quality control:
This reads the transcript and sorts every word by confidence score from lowest to highest. It shows you the 10 most uncertain words — the ones most likely to be transcription errors. A human reviewer can then jump to those timestamps in the original audio and fix them. This is much faster than proofreading the entire transcript from scratch!
import json
# Load the transcript (already downloaded from Object Storage)
with open("transcript.json", "r") as f:
transcript_data = json.load(f)
# Collect all words and their confidence scores
all_words = []
for token in transcript_data["transcriptions"][0]["tokens"]:
word = token.get("token", "")
# Skip punctuation — it doesn't have a meaningful confidence score
if word in (".", ",", "!", "?", ";", ":"):
continue
confidence = float(token.get("confidence", 1.0))
timestamp = token.get("startTime", "")
all_words.append({
"word": word,
"confidence": confidence,
"timestamp": timestamp
})
# Sort by confidence — lowest (most uncertain) first
all_words.sort(key=lambda x: x["confidence"])
# Print the 10 words the AI was least confident about
print("🔍 Top 10 words to manually verify (lowest confidence first):")
print(f"{'Rank'}<5 | {'Word':<20} | {'Confidence':<12} | {'Timestamp'}")
print("-" * 55)
for rank, word_info in enumerate(all_words[:10], start=1):
confidence_pct = word_info["confidence"] * 100
# Add a visual warning icon for very low confidence
warning = "⚠️ " if confidence_pct < 70 else " "
print(
f"{warning}{rank:<4} | "
f"{word_info['word']:<20} | "
f"{confidence_pct:<10.1f}% | "
f"{word_info['timestamp']}"
)
Example output:
🔍 Top 10 words to manually verify (lowest confidence first): Rank | Word | Confidence | Timestamp ------------------------------------------------------- ⚠️ 1 | Ranjodh | 58.2% | PT14.20S ⚠️ 2 | Krishnamurthy | 61.4% | PT45.80S 3 | approximately | 73.1% | PT18.30S 4 | synergies | 74.5% | PT67.10S 5 | paradigm | 75.8% | PT89.40S 6 | procurement | 77.2% | PT102.30S 7 | decentralised | 78.0% | PT134.50S 8 | bandwidth | 78.4% | PT156.20S 9 | tokenisation | 79.1% | PT178.60S 10 | regulatory | 79.9% | PT203.80S
Notice that proper names like "Ranjodh" and "Krishnamurthy" had the lowest confidence — that's expected! Names are harder for AI to guess because they don't appear often in training data. You can use OCI Speech's Custom Vocabulary feature to teach it your team's names and industry-specific words so that accuracy on those terms improves dramatically.
Real-World Use Cases — Where Is OCI Speech Used? 🏢
Let's look at real industries where OCI Speech is transforming work:
- 🏥 Healthcare — Doctors record patient consultations. OCI Speech transcribes them automatically, saving 2-3 hours of manual typing per day. The transcript goes directly into the Electronic Health Record system.
- 📞 Call Centers — Every customer call is automatically transcribed. Managers can search for specific complaint keywords across thousands of calls. OCI Language can then analyse sentiment on each call automatically!
- 📺 Media & Broadcasting — TV stations automatically generate subtitles for every program. SRT files created by OCI Speech are uploaded directly to the video platform. No manual captioning required.
- ⚖️ Legal — Court hearings are recorded and transcribed in real-time. Lawyers search the transcript for specific statements in seconds instead of scrubbing through hours of recordings.
- 🎓 Education — University lectures are automatically transcribed and captioned for accessibility. Students with hearing difficulties can follow every lecture in real-time using OCI's real-time speech service.
- 💰 Finance — Earnings calls, board meetings, and investor presentations are automatically transcribed and analysed by OCI Language for key mentions of metrics, risks, and forecasts.
Quick Summary 📝
What we learned about OCI Speech:
- OCI Speech → Converts audio and video files into accurate text using AI (ASR technology)
- Batch Transcription → Upload audio file to Object Storage → Create a Job → Get JSON/SRT output
- Real-Time Transcription → Live speech-to-text via WebSocket — words appear as you speak
- Oracle ASR Model → Best accuracy for English, Spanish, Portuguese
- Whisper Model → 50+ languages including Hindi, French, German, Japanese, and more
- Speaker Diarization → Identifies who said what — labels 2 to 16 speakers automatically
- Confidence Scores → Every word gets a 0-1 score — use it to target manual review effort
- Text Normalization → Converts spoken numbers, dates, addresses to proper written form
- Profanity Filtering → Mask, remove, or tag bad words automatically
- SRT Output → Ready-made subtitle files for video platforms
Happy building! 🎙️☁️
Comments
Post a Comment