Skip to main content

CoNLL vs JSONL in OCI Data Science: Dataset Formats for Machine Learning

Calculating read time…

Have you ever wondered how AI models actually learn to understand language?
How does a computer know that "Mumbai" is a city name and "Sachin" is a person's name?
The secret is hiding inside two special data formats — CoNLL and JSONL.

📂 Why Do Data Formats Matter in AI?

Before an AI model can learn anything, it needs data.
But data can be messy — names, sentences, labels all jumbled together.

Data formats are like filing systems — they organize your data so the AI can read it cleanly.
Wrong format = AI gets confused. Right format = AI learns perfectly! ✅

⚠️ Important:
OCI Data Science supports multiple data formats for NLP (Natural Language Processing) tasks.
The two most important ones for training language models are CoNLL and JSONL.

🗺️ The Big Picture — How Data Flows in OCI

Here is a simple flow showing how your data moves from raw text to a trained AI model in OCI:

📝
Raw Text
Data
→
🏷️
Label &
Annotate
→
📄
CoNLL or
JSONL Format
→
☁️
Upload to
OCI Storage
→
🤖
Train AI
Model (OCI)
→
🚀
Deploy &
Use Model

Simple, right? Now let's understand each format in detail! 🎯

📌 Part 1 — What is CoNLL Format?

CoNLL stands for Conference on Natural Language Learning.
It is a special text format invented to help researchers share and compare NLP datasets.

Think of CoNLL like a school attendance register.
Each student (word) gets their own row, and the teacher writes labels next to each name.

🔍 What Does CoNLL Look Like?

In CoNLL format, each word gets its own line.
Each line has the word and its label, separated by a space or tab.
A blank line means the sentence has ended.

Sachin   B-PER
Tendulkar   I-PER
plays   O
cricket   O
in   O
Mumbai   B-LOC
.   O

He   O
is   O
a   O
legend   O
.   O
🟢 Understanding the Labels:
  • B-PER → Beginning of a Person's name
  • I-PER → Inside (continuation) of a Person's name
  • B-LOC → Beginning of a Location name
  • O → Outside — just a regular word, no special label
This labeling system is called BIO Tagging (Beginning, Inside, Outside)!

🎒 Real-World Analogy for BIO Tagging

Imagine you are highlighting names in a book with coloured markers:

  • 🟡 Yellow marker on the FIRST word of a name = B- (Beginning)
  • 🟠 Orange marker on the NEXT words of the same name = I- (Inside)
  • ⬜ No marker on regular words = O (Outside)

So "Sachin Tendulkar" becomes → B-PER then I-PER.
And "Mumbai" becomes → B-LOC (it's just one word, so no I- needed).

🌟 Why is CoNLL Important in OCI Data Science?

  • ✅ Used for training Named Entity Recognition (NER) models — to find names, places, dates in text
  • ✅ Used for Part-of-Speech (POS) Tagging — knowing if a word is a noun, verb, adjective
  • ✅ Used for Dependency Parsing — understanding sentence grammar structure
  • ✅ OCI's Language AI Service and custom NLP models accept CoNLL-style annotated data
  • ✅ Most popular NLP libraries — spaCy, HuggingFace — can directly read CoNLL files

⌨️ CoNLL in Action — Python Code Examples

📋 What Will the Code Below Do?
This code will read a CoNLL file from your computer.
It will split it into sentences and print each word with its label — like reading that attendance register row by row!
Think of it as: "Open the register, read each student's name and their grade, print it on screen."
# Step 1: Read a CoNLL file and extract sentences with labels

def read_conll_file(filepath):
    sentences = []
    current_sentence = []

    with open(filepath, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()

            # Blank line = end of one sentence
            if line == "":
                if current_sentence:
                    sentences.append(current_sentence)
                    current_sentence = []
            else:
                # Each line has: word <tab> label
                parts = line.split()
                word = parts[0]
                label = parts[1] if len(parts) > 1 else 'O'
                current_sentence.append((word, label))

    # Don't forget the last sentence!
    if current_sentence:
        sentences.append(current_sentence)

    return sentences


# Let's test it!
sentences = read_conll_file('my_data.conll')

for i, sentence in enumerate(sentences):
    print(f"\n--- Sentence {i+1} ---")
    for word, label in sentence:
        print(f"  Word: {word:15s} | Label: {label}")

Output (example):

--- Sentence 1 ---
  Word: Sachin          | Label: B-PER
  Word: Tendulkar       | Label: I-PER
  Word: plays           | Label: O
  Word: cricket         | Label: O
  Word: in              | Label: O
  Word: Mumbai          | Label: B-LOC
  Word: .               | Label: O

See how organized that is? Every word has its label, neat and clean! 🎯

📋 What Will the Code Below Do?
This code will create a CoNLL file from scratch.
Imagine you have a sentence and you already know which words are names/places — this code writes that information into a proper CoNLL file so AI models can read it later!
# Step 2: Write your own CoNLL file from scratch

def write_conll_file(sentences, filepath):
    """
    sentences = list of lists
    Each inner list = [(word, label), (word, label), ...]
    """
    with open(filepath, 'w', encoding='utf-8') as f:
        for sentence in sentences:
            for word, label in sentence:
                f.write(f"{word}\t{label}\n")
            # Blank line separates sentences
            f.write("\n")

    print(f"✅ CoNLL file saved at: {filepath}")


# Example data — two sentences with BIO labels
my_sentences = [
    [
        ("Virat", "B-PER"),
        ("Kohli", "I-PER"),
        ("lives", "O"),
        ("in", "O"),
        ("Delhi", "B-LOC"),
        (".", "O")
    ],
    [
        ("Google", "B-ORG"),
        ("was", "O"),
        ("founded", "O"),
        ("in", "O"),
        ("California", "B-LOC"),
        (".", "O")
    ]
]

write_conll_file(my_sentences, 'my_output.conll')

Running this creates a clean CoNLL file ready for AI training! 🚀

🔬 Advanced CoNLL — Multi-Column Format

In real-world NLP projects, CoNLL files often have more than two columns.
For example, CoNLL-2003 (a famous dataset) has 4 columns:

Word       POS-Tag   Chunk-Tag   NER-Tag
-------    -------   ---------   -------
Sachin     NNP       B-NP        B-PER
Tendulkar  NNP       I-NP        I-PER
plays      VBZ       B-VP        O
cricket    NN        B-NP        O
in         IN        B-PP        O
Mumbai     NNP       B-NP        B-LOC
.          .         O           O
  • POS-Tag = Part of Speech (NNP = Proper Noun, VBZ = Verb, etc.)
  • Chunk-Tag = Groups of words (NP = Noun Phrase, VP = Verb Phrase)
  • NER-Tag = Named Entity (PER = Person, LOC = Location, ORG = Organisation)
💡 OCI Tip :
OCI Language AI Service now supports custom entity detection models trained on CoNLL-style data.
You can upload your annotated CoNLL dataset via OCI Object Storage and train a model
directly from the OCI Console — no deep ML expertise needed!

📌 Part 2 — What is JSONL Format?

JSONL stands for JSON Lines.
You might already know JSON — it's a popular format to store data like a dictionary.
JSONL is just JSON but with one JSON object per line.

🔍 JSON vs JSONL — What's the Difference?

📁 Regular JSON (List of objects)
[
  {"text": "I love OCI", "label": "positive"},
  {"text": "OCI is slow", "label": "negative"},
  {"text": "OCI is okay", "label": "neutral"}
]

❌ Must read the ENTIRE file to process even one record.

📄 JSONL (One JSON per line)
{"text": "I love OCI", "label": "positive"}
{"text": "OCI is slow", "label": "negative"}
{"text": "OCI is okay", "label": "neutral"}

✅ Read and process one line at a time — great for HUGE files!

🟢 Simple Analogy:
Regular JSON = A book — you need to open it fully to find anything.
JSONL = A set of index cards — each card is complete on its own, pick any card any time! 🃏

🌟 Why is JSONL Important in OCI Data Science?

  • ✅ Fine-tuning Large Language Models (LLMs) — OCI now supports LLM fine-tuning via JSONL prompt-response pairs
  • ✅ Text Classification datasets — sentiment analysis, spam detection, topic categorization
  • ✅ Question Answering datasets — training models to answer questions from passages
  • ✅ OCI Generative AI Service accepts JSONL for custom model fine-tuning
  • ✅ Streaming-friendly — you can process millions of records without loading everything into memory
  • ✅ HuggingFace's datasets library reads JSONL natively with one line of code

📐 JSONL Structures for Different NLP Tasks

Structure 1: Text Classification

{"text": "The new OCI GPU instance is blazing fast!", "label": "positive"}
{"text": "My OCI notebook keeps crashing today.", "label": "negative"}
{"text": "OCI pricing is competitive with AWS.", "label": "neutral"}

Structure 2: LLM Fine-Tuning (Chat Format — OCI Generative AI )

{"prompt": "What is OCI Object Storage?", 
"completion": "OCI Object Storage is a scalable cloud storage service by Oracle."}
{"prompt": "How do I create a notebook in OCI?",
"completion": "Go to OCI Console, navigate to Data Science, create a Project,
then create a Notebook Session."}
{"prompt": "What is CoNLL format?", "completion": "CoNLL is a text format where each word
in a sentence is on its own line with a label next to it."}

Structure 3: Named Entity Recognition (NER) in JSONL

{"text": "Sachin Tendulkar plays for Mumbai Indians.", 
"entities": [{"start": 0, "end": 16, "label": "PERSON"}, {"start": 27, "end": 42, "label": "ORG"}]} {"text": "Google opened a new office in Bengaluru.",
"entities": [{"start": 0, "end": 6, "label": "ORG"}, {"start": 30, "end": 39, "label": "LOCATION"}]}

Structure 4: Question Answering

{"context": "OCI was launched by Oracle in 2016.", 
"question": "When was OCI launched?", "answer": "2016"} {"context": "Python is a popular programming language created by Guido van Rossum.",
"question": "Who created Python?", "answer": "Guido van Rossum"}

⌨️ JSONL in Action — Python Code Examples

📋 What Will the Code Below Do?
This code will create a JSONL file for a text classification task.
Think of it as: "I have a list of customer reviews and I want to save them in a format that AI can use for training."
After running this code, you'll have a .jsonl file ready to upload to OCI!
import json

# Step 1: Create a JSONL file for text classification

reviews = [
    {"text": "OCI Data Science notebooks are super easy to use!", "label": "positive"},
    {"text": "Setting up OCI took a long time.", "label": "negative"},
    {"text": "OCI Object Storage works fine for our use case.", "label": "neutral"},
    {"text": "The OCI GPU clusters are incredibly fast for training!", "label": "positive"},
    {"text": "OCI documentation could be better for beginners.", "label": "negative"},
]

# Write each review as one JSON line
output_file = "oci_reviews.jsonl"

with open(output_file, 'w', encoding='utf-8') as f:
    for review in reviews:
        # json.dumps() converts dict to JSON string
        f.write(json.dumps(review) + '\n')

print(f"✅ JSONL file created: {output_file}")
print(f"   Total records: {len(reviews)}")

Output:

✅ JSONL file created: oci_reviews.jsonl
   Total records: 5
📋 What Will the Code Below Do?
This code will read a JSONL file and print each record.
It also counts how many "positive", "negative", and "neutral" labels exist — like a quick report card for your dataset!
import json
from collections import Counter

# Step 2: Read and analyze a JSONL file

def read_and_analyze_jsonl(filepath):
    records = []
    label_counts = Counter()

    with open(filepath, 'r', encoding='utf-8') as f:
        for line_number, line in enumerate(f, start=1):
            line = line.strip()
            if not line:
                continue  # Skip empty lines

            # Convert JSON string back to Python dictionary
            record = json.loads(line)
            records.append(record)
            label_counts[record.get('label', 'unknown')] += 1

            print(f"Record {line_number}: {record['text'][:50]}... → [{record['label']}]")

    print(f"\n📊 Dataset Summary:")
    print(f"   Total records: {len(records)}")
    for label, count in label_counts.items():
        print(f"   {label}: {count} records")

    return records

# Run the analysis
data = read_and_analyze_jsonl("oci_reviews.jsonl")

Output:

Record 1: OCI Data Science notebooks are super easy to use!... → [positive]
Record 2: Setting up OCI took a long time.... → [negative]
Record 3: OCI Object Storage works fine for our use case.... → [neutral]
Record 4: The OCI GPU clusters are incredibly fast for training!... → [positive]
Record 5: OCI documentation could be better for beginners.... → [negative]

📊 Dataset Summary:
   Total records: 5
   positive: 2 records
   negative: 2 records
   neutral: 1 records
📋 What Will the Code Below Do?
This code will create a JSONL file for LLM fine-tuning in OCI Generative AI 
You are teaching an AI to answer questions about OCI — like writing flashcards for the AI to study from!
import json

# Step 3: Create JSONL for OCI Generative AI Fine-Tuning
# OCI Generative AI uses "prompt" and "completion" fields

qa_pairs = [
    {
        "prompt": "What is OCI Data Science?",
        "completion": "OCI Data Science is a fully managed cloud platform by Oracle "
                      "that provides tools, compute, and storage for building and "
                      "deploying machine learning models."
    },
    {
        "prompt": "What is a Notebook Session in OCI?",
        "completion": "A Notebook Session in OCI is a Jupyter notebook environment "
                      "running on a cloud compute instance. You can write Python code, "
                      "train models, and visualize data directly in your browser."
    },
    {
        "prompt": "What is CoNLL format used for?",
        "completion": "CoNLL format is used to store annotated NLP data where each "
                      "word gets its own line with a label. It is commonly used for "
                      "NER, POS tagging, and dependency parsing tasks."
    },
    {
        "prompt": "How is JSONL different from JSON?",
        "completion": "JSON stores all data in a single array or object. JSONL stores "
                      "one JSON object per line, making it easy to stream and process "
                      "very large datasets without loading everything into memory."
    },
]

with open("oci_finetuning.jsonl", 'w', encoding='utf-8') as f:
    for pair in qa_pairs:
        f.write(json.dumps(pair) + '\n')

print(f"✅ Fine-tuning JSONL created with {len(qa_pairs)} Q&A pairs!")
print("   Ready to upload to OCI Object Storage for model fine-tuning.")

Perfect! Your fine-tuning dataset is ready for OCI Generative AI! 🤖

⚖️ CoNLL vs JSONL — Which One Should You Use?

This is one of the most common questions for beginners!
Here is a simple side-by-side comparison:

Feature CoNLL JSONL
Format Style One word per line, space/tab separated One JSON object per line
Best For NER, POS tagging, sequence labeling Text classification, LLM fine-tuning, QA
Human Readable? ✅ Very easy to read ✅ Easy with formatting
Handles Nested Data? ❌ No (flat structure only) ✅ Yes (full JSON nesting)
OCI Language AI ✅ Supported for NER training ✅ Supported for classification & fine-tuning
OCI GenAI Fine-Tuning ❌ Not typical ✅ Primary format
File Size Efficiency ⚠️ Can get large (one line per word) ✅ Compact per record
Library Support spaCy, HuggingFace, NLTK HuggingFace datasets, pandas, jsonlines
🟢 Simple Rule of Thumb:
👉 Use CoNLL when you need to label individual words in a sentence (NER, POS tagging).
👉 Use JSONL when you need to label entire texts or create prompt-response pairs for LLMs.

🛠️ End-to-End OCI Workflow — Using CoNLL & JSONL Together

In a real OCI AI project, you might actually use both formats!
Here's how a typical workflow looks:

1
Collect Raw Text Data
Gather your sentences, articles, customer reviews, medical notes — whatever your project needs.
2
Annotate for NER → Save as CoNLL
Use annotation tools (like Prodigy or Label Studio) to tag person/location/org names.
Export as CoNLL format.
3
Annotate for Classification → Save as JSONL
Assign sentiment/topic labels to full texts.
Export as JSONL format.
4
Upload Files to OCI Object Storage
Use OCI Console or Python SDK to upload your .conll and .jsonl files to a bucket.
5
Train Models in OCI Data Science Notebooks
Use HuggingFace + OCI GPU instances to train NER model (from CoNLL)
and classification model (from JSONL).
6
Deploy & Use via OCI Model Deployment
Deploy your trained models as REST API endpoints.
Your app can now extract entities AND classify text! 🎉

⌨️ Uploading Files to OCI Object Storage — Python Code

📋 What Will the Code Below Do?
This code will upload your CoNLL or JSONL file to OCI Object Storage.
Think of it like: "I want to save my data file in Oracle's cloud hard drive so the AI training notebook can access it."
You must have the OCI Python SDK installed: pip install oci
import oci

# Step 1: Set up OCI configuration
# (Make sure ~/.oci/config is set up with your credentials)
config = oci.config.from_file()

# Step 2: Create Object Storage client
object_storage = oci.object_storage.ObjectStorageClient(config)

# Step 3: Upload a file to OCI Object Storage

namespace = object_storage.get_namespace().data  # Your OCI tenancy namespace

bucket_name = "my-nlp-data-bucket"      # The bucket you created in OCI Console
local_file  = "my_data.conll"           # Your local CoNLL file
object_name = "datasets/my_data.conll"  # The path/name inside the bucket

with open(local_file, 'rb') as f:
    object_storage.put_object(
        namespace_name=namespace,
        bucket_name=bucket_name,
        object_name=object_name,
        put_object_body=f
    )

print(f"✅ Successfully uploaded '{local_file}' to OCI bucket '{bucket_name}'")
print(f"   Object path: {object_name}")

Your data file is now safely in the cloud and ready for AI training! ☁️

🤗 Using HuggingFace with CoNLL & JSONL in OCI (Best Practice)

The most popular way to train NLP models in OCI is by combining
HuggingFace Transformers with OCI's powerful GPU compute instances.

📋 What Will the Code Below Do?
This code will load a JSONL dataset using HuggingFace and prepare it for training.
It's like opening your JSONL file and handing it directly to an AI trainer who knows exactly what to do with it!
from datasets import load_dataset

# Step 1: Load JSONL file with HuggingFace datasets library
# This reads your JSONL file and creates a Dataset object

dataset = load_dataset(
    'json',
    data_files={
        'train': 'train_data.jsonl',
        'test':  'test_data.jsonl'
    }
)

print(dataset)
# Output: DatasetDict with 'train' and 'test' splits

print("\nFirst training example:")
print(dataset['train'][0])

# Output example:
# {'text': 'OCI Data Science notebooks are super easy to use!', 'label': 'positive'}
📋 What Will the Code Below Do?
This code will load a CoNLL-format NER dataset with HuggingFace.
HuggingFace has a special datasets loader for CoNLL files — no manual parsing needed!
Think of it as: "Let the library do the hard work of reading the file — I just call one function."
from datasets import load_dataset

# Step 2: Load CoNLL-format NER data using HuggingFace
# The 'conll2003' is a famous public NER dataset — great for learning!

dataset = load_dataset("conll2003")

# See what's inside
print(dataset)
print("\nFirst sentence tokens:", dataset['train'][0]['tokens'])
print("NER tags:             ", dataset['train'][0]['ner_tags'])

# Output:
# First sentence tokens: ['EU', 'rejects', 'German', 'call', ...]
# NER tags:              [3, 0, 7, 0, ...]
# (Numbers map to label names like B-ORG, O, B-LOC etc.)

# Check the label mapping
print("\nLabel names:", dataset['train'].features['ner_tags'].feature.names)
# Output: ['O', 'B-PER', 'I-PER', 'B-ORG', 'I-ORG', 'B-LOC', 'I-LOC', ...]
💡 OCI Tip:
When running HuggingFace training inside an OCI Notebook Session,
choose a VM.GPU.A10.1 or BM.GPU.H100.8 shape for faster model training.
OCI now offers NVIDIA H100 GPUs for enterprise-scale LLM fine-tuning! 🔥

🔄 Converting CoNLL to JSONL (and Back!)

Sometimes you need to convert between the two formats.
For example, you annotated data in CoNLL format but your model expects JSONL.

📋 What Will the Code Below Do?
This code will convert a CoNLL file to JSONL format.
It reads each sentence from CoNLL (word + label per line)
and saves it as a JSONL record with the full sentence text and entity list.
Like translating a document from one language to another! 🌐
import json

def conll_to_jsonl(conll_filepath, jsonl_filepath):
    """
    Converts a CoNLL file to JSONL format.
    Each sentence becomes one JSON record.
    """
    records = []
    current_tokens = []
    current_labels = []

    with open(conll_filepath, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()

            if line == "":
                # End of sentence — save as one JSONL record
                if current_tokens:
                    record = {
                        "tokens": current_tokens,
                        "labels": current_labels,
                        "text": " ".join(current_tokens)
                    }
                    records.append(record)
                    current_tokens = []
                    current_labels = []
            else:
                parts = line.split()
                current_tokens.append(parts[0])
                current_labels.append(parts[1] if len(parts) > 1 else 'O')

    # Handle last sentence if file doesn't end with blank line
    if current_tokens:
        records.append({
            "tokens": current_tokens,
            "labels": current_labels,
            "text": " ".join(current_tokens)
        })

    # Write all records to JSONL
    with open(jsonl_filepath, 'w', encoding='utf-8') as out:
        for record in records:
            out.write(json.dumps(record) + '\n')

    print(f"✅ Converted {len(records)} sentences from CoNLL to JSONL")
    print(f"   Output: {jsonl_filepath}")

# Run the conversion
conll_to_jsonl("my_data.conll", "my_data.jsonl")

Example JSONL Output (one line per sentence):

{"tokens": ["Sachin", "Tendulkar", "plays", "in", "Mumbai", "."], 
"labels": ["B-PER", "I-PER", "O", "O", "B-LOC", "O"], "text": "Sachin Tendulkar plays in Mumbai ."}

✅ Best Practices — DOs and DON'Ts

🟢 DOs — Follow These Always!
  • ✅ Always use UTF-8 encoding when reading/writing CoNLL and JSONL files
  • ✅ Always end CoNLL sentences with a blank line
  • ✅ Always validate your JSONL — each line must be a valid JSON object
  • ✅ Use train/validation/test splits — typically 70% / 15% / 15%
  • ✅ Store your data in OCI Object Storage (not local disk) for sharing with notebooks
  • ✅ Use consistent label names throughout your CoNLL file (B-PER, not sometimes PER_B)
  • ✅ For LLM fine-tuning JSONL — keep prompt/completion pairs concise and focused
🔴 DON'Ts — Avoid These Mistakes!
  • ❌ Don't mix label formats (don't use both BIO and IO tagging in the same file)
  • ❌ Don't store entire datasets as a single JSON array — use JSONL for large files
  • ❌ Don't skip the blank line between sentences in CoNLL — it will break parsers
  • ❌ Don't hardcode file paths — use OCI Object Storage URIs for portability
  • ❌ Don't train on unbalanced datasets without checking label distribution first
  • ❌ Don't forget to validate your CoNLL file — a single malformed line can crash training
⚠️ Important Warning:
When uploading data to OCI Object Storage for AI training — never include personally identifiable information (PII)
like real names, phone numbers, or addresses unless you have proper data governance approvals in place.
OCI provides Data Safe and OCI Vault tools to help you protect sensitive data!

⌨️ Quick Validation — Check Your Files Before Training

📋 What Will the Code Below Do?
This code will validate your JSONL file — it checks every line to make sure it's proper JSON.
It also checks that required fields like "text" and "label" are present.
Like a spell-checker, but for your data file — catching errors before they break your training! 🔍
import json

def validate_jsonl(filepath, required_keys=None):
    """
    Validates a JSONL file.
    Checks that each line is valid JSON and has required keys.
    """
    if required_keys is None:
        required_keys = []

    errors = []
    valid_count = 0

    with open(filepath, 'r', encoding='utf-8') as f:
        for line_num, line in enumerate(f, start=1):
            line = line.strip()
            if not line:
                continue

            try:
                record = json.loads(line)
                # Check for required keys
                for key in required_keys:
                    if key not in record:
                        errors.append(f"Line {line_num}: Missing key '{key}'")
                        continue
                valid_count += 1
            except json.JSONDecodeError as e:
                errors.append(f"Line {line_num}: Invalid JSON — {e}")

    print(f"\n📋 Validation Report for: {filepath}")
    print(f"   ✅ Valid records:  {valid_count}")
    print(f"   ❌ Errors found:   {len(errors)}")

    if errors:
        print("\n🔴 Error Details:")
        for err in errors[:10]:  # Show first 10 errors
            print(f"   {err}")
    else:
        print("\n🎉 File is perfectly valid! Ready for OCI training.")


# Validate your classification JSONL
validate_jsonl("oci_reviews.jsonl", required_keys=["text", "label"])

Output (if valid):

📋 Validation Report for: oci_reviews.jsonl
   ✅ Valid records:  5
   ❌ Errors found:   0

🎉 File is perfectly valid! Ready for OCI training.

📝 Quick Summary — What We Learned

  • CoNLL Format → One word per line with a label, blank line between sentences.
    Used for NER, POS tagging, sequence labeling tasks.
  • BIO Tagging → B = Beginning of an entity, I = Inside an entity, O = Outside (regular word).
  • JSONL Format → One JSON object per line. Used for text classification, LLM fine-tuning, QA datasets.
  • OCI Data Science → Cloud platform to store (Object Storage), process (Notebooks), train, and deploy AI models.
  • When to use CoNLL → When labeling individual words in a sentence.
  • When to use JSONL → When labeling full texts or creating prompt-response training pairs for LLMs.
  • Conversion → You can convert CoNLL → JSONL using Python when your model requires a different format.
  • HuggingFace + OCI → The #1 combo for training NLP models on Oracle Cloud.

Comments