Have you ever wondered how AI models actually learn to understand language?
How does a computer know that "Mumbai" is a city name and "Sachin" is a person's name?
The secret is hiding inside two special data formats — CoNLL and JSONL.
📂 Why Do Data Formats Matter in AI?
Before an AI model can learn anything, it needs data.
But data can be messy — names, sentences, labels all jumbled together.
Data formats are like filing systems — they organize your data so the AI can read it cleanly.
Wrong format = AI gets confused. Right format = AI learns perfectly! ✅
OCI Data Science supports multiple data formats for NLP (Natural Language Processing) tasks.
The two most important ones for training language models are CoNLL and JSONL.
🗺️ The Big Picture — How Data Flows in OCI
Here is a simple flow showing how your data moves from raw text to a trained AI model in OCI:
Raw Text
Data
Label &
Annotate
CoNLL or
JSONL Format
Upload to
OCI Storage
Train AI
Model (OCI)
Deploy &
Use Model
Simple, right? Now let's understand each format in detail! 🎯
📌 Part 1 — What is CoNLL Format?
CoNLL stands for Conference on Natural Language Learning.
It is a special text format invented to help researchers share and compare NLP datasets.
Think of CoNLL like a school attendance register.
Each student (word) gets their own row, and the teacher writes labels next to each name.
🔍 What Does CoNLL Look Like?
In CoNLL format, each word gets its own line.
Each line has the word and its label, separated by a space or tab.
A blank line means the sentence has ended.
Sachin B-PER Tendulkar I-PER plays O cricket O in O Mumbai B-LOC . O He O is O a O legend O . O
- B-PER → Beginning of a Person's name
- I-PER → Inside (continuation) of a Person's name
- B-LOC → Beginning of a Location name
- O → Outside — just a regular word, no special label
🎒 Real-World Analogy for BIO Tagging
Imagine you are highlighting names in a book with coloured markers:
- 🟡 Yellow marker on the FIRST word of a name = B- (Beginning)
- 🟠 Orange marker on the NEXT words of the same name = I- (Inside)
- ⬜ No marker on regular words = O (Outside)
So "Sachin Tendulkar" becomes → B-PER then I-PER.
And "Mumbai" becomes → B-LOC (it's just one word, so no I- needed).
🌟 Why is CoNLL Important in OCI Data Science?
- ✅ Used for training Named Entity Recognition (NER) models — to find names, places, dates in text
- ✅ Used for Part-of-Speech (POS) Tagging — knowing if a word is a noun, verb, adjective
- ✅ Used for Dependency Parsing — understanding sentence grammar structure
- ✅ OCI's Language AI Service and custom NLP models accept CoNLL-style annotated data
- ✅ Most popular NLP libraries — spaCy, HuggingFace — can directly read CoNLL files
⌨️ CoNLL in Action — Python Code Examples
This code will read a CoNLL file from your computer.
It will split it into sentences and print each word with its label — like reading that attendance register row by row!
Think of it as: "Open the register, read each student's name and their grade, print it on screen."
# Step 1: Read a CoNLL file and extract sentences with labels
def read_conll_file(filepath):
sentences = []
current_sentence = []
with open(filepath, 'r', encoding='utf-8') as f:
for line in f:
line = line.strip()
# Blank line = end of one sentence
if line == "":
if current_sentence:
sentences.append(current_sentence)
current_sentence = []
else:
# Each line has: word <tab> label
parts = line.split()
word = parts[0]
label = parts[1] if len(parts) > 1 else 'O'
current_sentence.append((word, label))
# Don't forget the last sentence!
if current_sentence:
sentences.append(current_sentence)
return sentences
# Let's test it!
sentences = read_conll_file('my_data.conll')
for i, sentence in enumerate(sentences):
print(f"\n--- Sentence {i+1} ---")
for word, label in sentence:
print(f" Word: {word:15s} | Label: {label}")
Output (example):
--- Sentence 1 --- Word: Sachin | Label: B-PER Word: Tendulkar | Label: I-PER Word: plays | Label: O Word: cricket | Label: O Word: in | Label: O Word: Mumbai | Label: B-LOC Word: . | Label: O
See how organized that is? Every word has its label, neat and clean! 🎯
This code will create a CoNLL file from scratch.
Imagine you have a sentence and you already know which words are names/places — this code writes that information into a proper CoNLL file so AI models can read it later!
# Step 2: Write your own CoNLL file from scratch
def write_conll_file(sentences, filepath):
"""
sentences = list of lists
Each inner list = [(word, label), (word, label), ...]
"""
with open(filepath, 'w', encoding='utf-8') as f:
for sentence in sentences:
for word, label in sentence:
f.write(f"{word}\t{label}\n")
# Blank line separates sentences
f.write("\n")
print(f"✅ CoNLL file saved at: {filepath}")
# Example data — two sentences with BIO labels
my_sentences = [
[
("Virat", "B-PER"),
("Kohli", "I-PER"),
("lives", "O"),
("in", "O"),
("Delhi", "B-LOC"),
(".", "O")
],
[
("Google", "B-ORG"),
("was", "O"),
("founded", "O"),
("in", "O"),
("California", "B-LOC"),
(".", "O")
]
]
write_conll_file(my_sentences, 'my_output.conll')
Running this creates a clean CoNLL file ready for AI training! 🚀
🔬 Advanced CoNLL — Multi-Column Format
In real-world NLP projects, CoNLL files often have more than two columns.
For example, CoNLL-2003 (a famous dataset) has 4 columns:
Word POS-Tag Chunk-Tag NER-Tag ------- ------- --------- ------- Sachin NNP B-NP B-PER Tendulkar NNP I-NP I-PER plays VBZ B-VP O cricket NN B-NP O in IN B-PP O Mumbai NNP B-NP B-LOC . . O O
- POS-Tag = Part of Speech (NNP = Proper Noun, VBZ = Verb, etc.)
- Chunk-Tag = Groups of words (NP = Noun Phrase, VP = Verb Phrase)
- NER-Tag = Named Entity (PER = Person, LOC = Location, ORG = Organisation)
OCI Language AI Service now supports custom entity detection models trained on CoNLL-style data.
You can upload your annotated CoNLL dataset via OCI Object Storage and train a model
directly from the OCI Console — no deep ML expertise needed!
📌 Part 2 — What is JSONL Format?
JSONL stands for JSON Lines.
You might already know JSON — it's a popular format to store data like a dictionary.
JSONL is just JSON but with one JSON object per line.
🔍 JSON vs JSONL — What's the Difference?
[
{"text": "I love OCI", "label": "positive"},
{"text": "OCI is slow", "label": "negative"},
{"text": "OCI is okay", "label": "neutral"}
]
❌ Must read the ENTIRE file to process even one record.
{"text": "I love OCI", "label": "positive"}
{"text": "OCI is slow", "label": "negative"}
{"text": "OCI is okay", "label": "neutral"}
✅ Read and process one line at a time — great for HUGE files!
Regular JSON = A book — you need to open it fully to find anything.
JSONL = A set of index cards — each card is complete on its own, pick any card any time! 🃏
🌟 Why is JSONL Important in OCI Data Science?
- ✅ Fine-tuning Large Language Models (LLMs) — OCI now supports LLM fine-tuning via JSONL prompt-response pairs
- ✅ Text Classification datasets — sentiment analysis, spam detection, topic categorization
- ✅ Question Answering datasets — training models to answer questions from passages
- ✅ OCI Generative AI Service accepts JSONL for custom model fine-tuning
- ✅ Streaming-friendly — you can process millions of records without loading everything into memory
- ✅ HuggingFace's
datasetslibrary reads JSONL natively with one line of code
📐 JSONL Structures for Different NLP Tasks
Structure 1: Text Classification
{"text": "The new OCI GPU instance is blazing fast!", "label": "positive"}
{"text": "My OCI notebook keeps crashing today.", "label": "negative"}
{"text": "OCI pricing is competitive with AWS.", "label": "neutral"}
Structure 2: LLM Fine-Tuning (Chat Format — OCI Generative AI )
{"prompt": "What is OCI Object Storage?",
"completion": "OCI Object Storage is a scalable cloud storage service by Oracle."}
{"prompt": "How do I create a notebook in OCI?",
"completion": "Go to OCI Console, navigate to Data Science, create a Project,
then create a Notebook Session."}
{"prompt": "What is CoNLL format?", "completion": "CoNLL is a text format where each word
in a sentence is on its own line with a label next to it."}
Structure 3: Named Entity Recognition (NER) in JSONL
{"text": "Sachin Tendulkar plays for Mumbai Indians.",
"entities": [{"start": 0, "end": 16, "label": "PERSON"}, {"start": 27, "end": 42, "label": "ORG"}]}
{"text": "Google opened a new office in Bengaluru.",
"entities": [{"start": 0, "end": 6, "label": "ORG"}, {"start": 30, "end": 39, "label": "LOCATION"}]}
Structure 4: Question Answering
{"context": "OCI was launched by Oracle in 2016.",
"question": "When was OCI launched?", "answer": "2016"}
{"context": "Python is a popular programming language created by Guido van Rossum.",
"question": "Who created Python?", "answer": "Guido van Rossum"}
⌨️ JSONL in Action — Python Code Examples
This code will create a JSONL file for a text classification task.
Think of it as: "I have a list of customer reviews and I want to save them in a format that AI can use for training."
After running this code, you'll have a
.jsonl file ready to upload to OCI!
import json
# Step 1: Create a JSONL file for text classification
reviews = [
{"text": "OCI Data Science notebooks are super easy to use!", "label": "positive"},
{"text": "Setting up OCI took a long time.", "label": "negative"},
{"text": "OCI Object Storage works fine for our use case.", "label": "neutral"},
{"text": "The OCI GPU clusters are incredibly fast for training!", "label": "positive"},
{"text": "OCI documentation could be better for beginners.", "label": "negative"},
]
# Write each review as one JSON line
output_file = "oci_reviews.jsonl"
with open(output_file, 'w', encoding='utf-8') as f:
for review in reviews:
# json.dumps() converts dict to JSON string
f.write(json.dumps(review) + '\n')
print(f"✅ JSONL file created: {output_file}")
print(f" Total records: {len(reviews)}")
Output:
✅ JSONL file created: oci_reviews.jsonl Total records: 5
This code will read a JSONL file and print each record.
It also counts how many "positive", "negative", and "neutral" labels exist — like a quick report card for your dataset!
import json
from collections import Counter
# Step 2: Read and analyze a JSONL file
def read_and_analyze_jsonl(filepath):
records = []
label_counts = Counter()
with open(filepath, 'r', encoding='utf-8') as f:
for line_number, line in enumerate(f, start=1):
line = line.strip()
if not line:
continue # Skip empty lines
# Convert JSON string back to Python dictionary
record = json.loads(line)
records.append(record)
label_counts[record.get('label', 'unknown')] += 1
print(f"Record {line_number}: {record['text'][:50]}... → [{record['label']}]")
print(f"\n📊 Dataset Summary:")
print(f" Total records: {len(records)}")
for label, count in label_counts.items():
print(f" {label}: {count} records")
return records
# Run the analysis
data = read_and_analyze_jsonl("oci_reviews.jsonl")
Output:
Record 1: OCI Data Science notebooks are super easy to use!... → [positive] Record 2: Setting up OCI took a long time.... → [negative] Record 3: OCI Object Storage works fine for our use case.... → [neutral] Record 4: The OCI GPU clusters are incredibly fast for training!... → [positive] Record 5: OCI documentation could be better for beginners.... → [negative] 📊 Dataset Summary: Total records: 5 positive: 2 records negative: 2 records neutral: 1 records
This code will create a JSONL file for LLM fine-tuning in OCI Generative AI
You are teaching an AI to answer questions about OCI — like writing flashcards for the AI to study from!
import json
# Step 3: Create JSONL for OCI Generative AI Fine-Tuning
# OCI Generative AI uses "prompt" and "completion" fields
qa_pairs = [
{
"prompt": "What is OCI Data Science?",
"completion": "OCI Data Science is a fully managed cloud platform by Oracle "
"that provides tools, compute, and storage for building and "
"deploying machine learning models."
},
{
"prompt": "What is a Notebook Session in OCI?",
"completion": "A Notebook Session in OCI is a Jupyter notebook environment "
"running on a cloud compute instance. You can write Python code, "
"train models, and visualize data directly in your browser."
},
{
"prompt": "What is CoNLL format used for?",
"completion": "CoNLL format is used to store annotated NLP data where each "
"word gets its own line with a label. It is commonly used for "
"NER, POS tagging, and dependency parsing tasks."
},
{
"prompt": "How is JSONL different from JSON?",
"completion": "JSON stores all data in a single array or object. JSONL stores "
"one JSON object per line, making it easy to stream and process "
"very large datasets without loading everything into memory."
},
]
with open("oci_finetuning.jsonl", 'w', encoding='utf-8') as f:
for pair in qa_pairs:
f.write(json.dumps(pair) + '\n')
print(f"✅ Fine-tuning JSONL created with {len(qa_pairs)} Q&A pairs!")
print(" Ready to upload to OCI Object Storage for model fine-tuning.")
Perfect! Your fine-tuning dataset is ready for OCI Generative AI! 🤖
⚖️ CoNLL vs JSONL — Which One Should You Use?
This is one of the most common questions for beginners!
Here is a simple side-by-side comparison:
| Feature | CoNLL | JSONL |
|---|---|---|
| Format Style | One word per line, space/tab separated | One JSON object per line |
| Best For | NER, POS tagging, sequence labeling | Text classification, LLM fine-tuning, QA |
| Human Readable? | ✅ Very easy to read | ✅ Easy with formatting |
| Handles Nested Data? | ❌ No (flat structure only) | ✅ Yes (full JSON nesting) |
| OCI Language AI | ✅ Supported for NER training | ✅ Supported for classification & fine-tuning |
| OCI GenAI Fine-Tuning | ❌ Not typical | ✅ Primary format |
| File Size Efficiency | ⚠️ Can get large (one line per word) | ✅ Compact per record |
| Library Support | spaCy, HuggingFace, NLTK | HuggingFace datasets, pandas, jsonlines |
👉 Use CoNLL when you need to label individual words in a sentence (NER, POS tagging).
👉 Use JSONL when you need to label entire texts or create prompt-response pairs for LLMs.
🛠️ End-to-End OCI Workflow — Using CoNLL & JSONL Together
In a real OCI AI project, you might actually use both formats!
Here's how a typical workflow looks:
Gather your sentences, articles, customer reviews, medical notes — whatever your project needs.
Use annotation tools (like Prodigy or Label Studio) to tag person/location/org names.
Export as CoNLL format.
Assign sentiment/topic labels to full texts.
Export as JSONL format.
Use OCI Console or Python SDK to upload your
.conll and .jsonl files to a bucket.
Use HuggingFace + OCI GPU instances to train NER model (from CoNLL)
and classification model (from JSONL).
Deploy your trained models as REST API endpoints.
Your app can now extract entities AND classify text! 🎉
⌨️ Uploading Files to OCI Object Storage — Python Code
This code will upload your CoNLL or JSONL file to OCI Object Storage.
Think of it like: "I want to save my data file in Oracle's cloud hard drive so the AI training notebook can access it."
You must have the OCI Python SDK installed:
pip install oci
import oci
# Step 1: Set up OCI configuration
# (Make sure ~/.oci/config is set up with your credentials)
config = oci.config.from_file()
# Step 2: Create Object Storage client
object_storage = oci.object_storage.ObjectStorageClient(config)
# Step 3: Upload a file to OCI Object Storage
namespace = object_storage.get_namespace().data # Your OCI tenancy namespace
bucket_name = "my-nlp-data-bucket" # The bucket you created in OCI Console
local_file = "my_data.conll" # Your local CoNLL file
object_name = "datasets/my_data.conll" # The path/name inside the bucket
with open(local_file, 'rb') as f:
object_storage.put_object(
namespace_name=namespace,
bucket_name=bucket_name,
object_name=object_name,
put_object_body=f
)
print(f"✅ Successfully uploaded '{local_file}' to OCI bucket '{bucket_name}'")
print(f" Object path: {object_name}")
Your data file is now safely in the cloud and ready for AI training! ☁️
🤗 Using HuggingFace with CoNLL & JSONL in OCI (Best Practice)
The most popular way to train NLP models in OCI is by combining
HuggingFace Transformers with OCI's powerful GPU compute instances.
This code will load a JSONL dataset using HuggingFace and prepare it for training.
It's like opening your JSONL file and handing it directly to an AI trainer who knows exactly what to do with it!
from datasets import load_dataset
# Step 1: Load JSONL file with HuggingFace datasets library
# This reads your JSONL file and creates a Dataset object
dataset = load_dataset(
'json',
data_files={
'train': 'train_data.jsonl',
'test': 'test_data.jsonl'
}
)
print(dataset)
# Output: DatasetDict with 'train' and 'test' splits
print("\nFirst training example:")
print(dataset['train'][0])
# Output example:
# {'text': 'OCI Data Science notebooks are super easy to use!', 'label': 'positive'}
This code will load a CoNLL-format NER dataset with HuggingFace.
HuggingFace has a special
datasets loader for CoNLL files — no manual parsing needed!Think of it as: "Let the library do the hard work of reading the file — I just call one function."
from datasets import load_dataset
# Step 2: Load CoNLL-format NER data using HuggingFace
# The 'conll2003' is a famous public NER dataset — great for learning!
dataset = load_dataset("conll2003")
# See what's inside
print(dataset)
print("\nFirst sentence tokens:", dataset['train'][0]['tokens'])
print("NER tags: ", dataset['train'][0]['ner_tags'])
# Output:
# First sentence tokens: ['EU', 'rejects', 'German', 'call', ...]
# NER tags: [3, 0, 7, 0, ...]
# (Numbers map to label names like B-ORG, O, B-LOC etc.)
# Check the label mapping
print("\nLabel names:", dataset['train'].features['ner_tags'].feature.names)
# Output: ['O', 'B-PER', 'I-PER', 'B-ORG', 'I-ORG', 'B-LOC', 'I-LOC', ...]
When running HuggingFace training inside an OCI Notebook Session,
choose a VM.GPU.A10.1 or BM.GPU.H100.8 shape for faster model training.
OCI now offers NVIDIA H100 GPUs for enterprise-scale LLM fine-tuning! 🔥
🔄 Converting CoNLL to JSONL (and Back!)
Sometimes you need to convert between the two formats.
For example, you annotated data in CoNLL format but your model expects JSONL.
This code will convert a CoNLL file to JSONL format.
It reads each sentence from CoNLL (word + label per line)
and saves it as a JSONL record with the full sentence text and entity list.
Like translating a document from one language to another! 🌐
import json
def conll_to_jsonl(conll_filepath, jsonl_filepath):
"""
Converts a CoNLL file to JSONL format.
Each sentence becomes one JSON record.
"""
records = []
current_tokens = []
current_labels = []
with open(conll_filepath, 'r', encoding='utf-8') as f:
for line in f:
line = line.strip()
if line == "":
# End of sentence — save as one JSONL record
if current_tokens:
record = {
"tokens": current_tokens,
"labels": current_labels,
"text": " ".join(current_tokens)
}
records.append(record)
current_tokens = []
current_labels = []
else:
parts = line.split()
current_tokens.append(parts[0])
current_labels.append(parts[1] if len(parts) > 1 else 'O')
# Handle last sentence if file doesn't end with blank line
if current_tokens:
records.append({
"tokens": current_tokens,
"labels": current_labels,
"text": " ".join(current_tokens)
})
# Write all records to JSONL
with open(jsonl_filepath, 'w', encoding='utf-8') as out:
for record in records:
out.write(json.dumps(record) + '\n')
print(f"✅ Converted {len(records)} sentences from CoNLL to JSONL")
print(f" Output: {jsonl_filepath}")
# Run the conversion
conll_to_jsonl("my_data.conll", "my_data.jsonl")
Example JSONL Output (one line per sentence):
{"tokens": ["Sachin", "Tendulkar", "plays", "in", "Mumbai", "."],
"labels": ["B-PER", "I-PER", "O", "O", "B-LOC", "O"], "text": "Sachin Tendulkar plays in Mumbai ."}
✅ Best Practices — DOs and DON'Ts
- ✅ Always use UTF-8 encoding when reading/writing CoNLL and JSONL files
- ✅ Always end CoNLL sentences with a blank line
- ✅ Always validate your JSONL — each line must be a valid JSON object
- ✅ Use train/validation/test splits — typically 70% / 15% / 15%
- ✅ Store your data in OCI Object Storage (not local disk) for sharing with notebooks
- ✅ Use consistent label names throughout your CoNLL file (B-PER, not sometimes PER_B)
- ✅ For LLM fine-tuning JSONL — keep prompt/completion pairs concise and focused
- ❌ Don't mix label formats (don't use both BIO and IO tagging in the same file)
- ❌ Don't store entire datasets as a single JSON array — use JSONL for large files
- ❌ Don't skip the blank line between sentences in CoNLL — it will break parsers
- ❌ Don't hardcode file paths — use OCI Object Storage URIs for portability
- ❌ Don't train on unbalanced datasets without checking label distribution first
- ❌ Don't forget to validate your CoNLL file — a single malformed line can crash training
When uploading data to OCI Object Storage for AI training — never include personally identifiable information (PII)
like real names, phone numbers, or addresses unless you have proper data governance approvals in place.
OCI provides Data Safe and OCI Vault tools to help you protect sensitive data!
⌨️ Quick Validation — Check Your Files Before Training
This code will validate your JSONL file — it checks every line to make sure it's proper JSON.
It also checks that required fields like "text" and "label" are present.
Like a spell-checker, but for your data file — catching errors before they break your training! 🔍
import json
def validate_jsonl(filepath, required_keys=None):
"""
Validates a JSONL file.
Checks that each line is valid JSON and has required keys.
"""
if required_keys is None:
required_keys = []
errors = []
valid_count = 0
with open(filepath, 'r', encoding='utf-8') as f:
for line_num, line in enumerate(f, start=1):
line = line.strip()
if not line:
continue
try:
record = json.loads(line)
# Check for required keys
for key in required_keys:
if key not in record:
errors.append(f"Line {line_num}: Missing key '{key}'")
continue
valid_count += 1
except json.JSONDecodeError as e:
errors.append(f"Line {line_num}: Invalid JSON — {e}")
print(f"\n📋 Validation Report for: {filepath}")
print(f" ✅ Valid records: {valid_count}")
print(f" ❌ Errors found: {len(errors)}")
if errors:
print("\n🔴 Error Details:")
for err in errors[:10]: # Show first 10 errors
print(f" {err}")
else:
print("\n🎉 File is perfectly valid! Ready for OCI training.")
# Validate your classification JSONL
validate_jsonl("oci_reviews.jsonl", required_keys=["text", "label"])
Output (if valid):
📋 Validation Report for: oci_reviews.jsonl ✅ Valid records: 5 ❌ Errors found: 0 🎉 File is perfectly valid! Ready for OCI training.
📝 Quick Summary — What We Learned
-
CoNLL Format →
One word per line with a label, blank line between sentences.
Used for NER, POS tagging, sequence labeling tasks. - BIO Tagging → B = Beginning of an entity, I = Inside an entity, O = Outside (regular word).
- JSONL Format → One JSON object per line. Used for text classification, LLM fine-tuning, QA datasets.
- OCI Data Science → Cloud platform to store (Object Storage), process (Notebooks), train, and deploy AI models.
- When to use CoNLL → When labeling individual words in a sentence.
- When to use JSONL → When labeling full texts or creating prompt-response training pairs for LLMs.
- Conversion → You can convert CoNLL → JSONL using Python when your model requires a different format.
- HuggingFace + OCI → The #1 combo for training NLP models on Oracle Cloud.
Comments
Post a Comment