Think of Vector Embeddings like converting words into GPS coordinates. Just as GPS uses numbers (latitude, longitude) to represent locations, embeddings use numbers to represent meaning! Words with similar meanings get similar coordinates.
What Are Vector Embeddings?
Vector embeddings are numerical representations of words, sentences, or any text. Instead of treating words as simple labels, we convert them into lists of numbers that capture their meaning.
- Input: Text like "king", "queen", "cat"
- Process: Convert to numerical vectors
- Output: Arrays of numbers like [0.2, -0.5, 0.8, ...]
Why Do We Need Embeddings?
The Problem: Computers Don't Understand Words
Computers only understand numbers. How do we make them understand language?
Bad Approach - One-Hot Encoding:
cat = [1, 0, 0, 0, 0]
dog = [0, 1, 0, 0, 0]
king = [0, 0, 1, 0, 0]
queen = [0, 0, 0, 1, 0]
tiger = [0, 0, 0, 0, 1]
Problems:
- All words are equally distant (no similarity captured)
- "cat" and "dog" should be similar (both animals) but they're not!
- Huge vectors (need one position for every word in vocabulary)
- Millions of words = millions of dimensions!
Good Approach - Dense Embeddings:
cat = [0.8, 0.3, -0.1, 0.5, ...] # 768 dimensions
dog = [0.7, 0.4, -0.2, 0.6, ...] # Similar to cat!
king = [-0.2, 0.9, 0.3, -0.4, ...]
queen = [-0.1, 0.8, 0.4, -0.3, ...] # Similar to king!
tiger = [0.6, 0.4, -0.3, 0.7, ...] # Similar to cat & dog!
Benefits:
- Fixed size (typically 768 or 1024 dimensions)
- Similar words have similar vectors
- Can compute similarity using math!
- Captures semantic relationships
Real Example: Semantic Similarity
Let's see how embeddings capture meaning with actual numbers!
We measure how similar two embeddings are using cosine similarity (ranges from -1 to 1):
- 1.0: Identical meaning
- 0.9: Very similar
- 0.5: Somewhat related
- 0.0: Unrelated
- -1.0: Opposite meaning
| Word Pair | Similarity Score | Interpretation |
|---|---|---|
| "king" ↔ "queen" | 0.85 | Very similar (both royalty) |
| "cat" ↔ "dog" | 0.80 | Very similar (both pets) |
| "king" ↔ "man" | 0.65 | Related (both male) |
| "cat" ↔ "tiger" | 0.75 | Similar (both felines) |
| "cat" ↔ "car" | 0.15 | Barely related |
| "hot" ↔ "cold" | -0.40 | Opposite meanings! |
The Famous King - Queen Example
Embeddings capture amazing semantic relationships! Here's the most famous example:
king - man + woman ≈ queen
What this means:
- Take the embedding for "king"
- Subtract the embedding for "man" (removes male aspect)
- Add the embedding for "woman" (adds female aspect)
- Result is very close to the embedding for "queen"!
More examples that work:
Paris - France + Italy ≈ Rome
walked - walk + run ≈ ran
bigger - big + small ≈ smaller
How Are Embeddings Created?
Method 1: Word2Vec (Classic Approach)
Word2Vec learns embeddings by predicting words from context.
Example Training:
Sentence: "The cat sat on the mat"
- Task: Given "cat", predict nearby words: "the", "sat", "on"
- Training: Adjust embeddings so similar contexts create similar vectors
- Result: "cat" and "dog" get similar embeddings (appear in similar contexts)
# Simplified Word2Vec concept
"The cat sat on the mat"
"The dog sat on the mat"
# Both "cat" and "dog" appear in same context:
# [The, ___, sat, on, the, mat]
# So they get similar embeddings!
Method 2: Transformer-based (Modern LLMs)
Modern models like BERT, GPT create contextual embeddings.
- Word2Vec (Static): "bank" always has same embedding
- BERT/GPT (Contextual): "bank" has different embeddings based on context!
Example - Word "bank":
| Sentence | Context | Embedding Type |
|---|---|---|
| "I went to the bank to deposit money" | Financial institution | Embedding A |
| "We sat by the river bank" | Land beside water | Embedding B |
With contextual embeddings, these two "bank" instances get DIFFERENT vectors! 🎯
Embeddings in Modern LLMs (GPT, Claude, etc.)
Step-by-Step: From Text to Understanding
Let's trace what happens when you give a prompt to an LLM:
Your Prompt: "Write a poem about cats"
Input: "Write a poem about cats"
Tokens: ["Write", "a", "poem", "about", "cats"]
Token IDs: [1234, 56, 789, 234, 456]
Token ID 1234 (Write) → Embedding: [0.2, -0.5, 0.8, ..., 0.3] (768 dims)
Token ID 56 (a) → Embedding: [-0.1, 0.3, -0.2, ..., 0.7]
Token ID 789 (poem) → Embedding: [0.5, 0.6, -0.3, ..., -0.1]
Token ID 234 (about) → Embedding: [0.1, -0.2, 0.4, ..., 0.2]
Token ID 456 (cats) → Embedding: [0.8, 0.3, -0.1, ..., 0.5]
# Position matters! "cats eat fish" ≠ "fish eat cats"
Embedding_final = Embedding_token + Embedding_position
# Each layer refines the embeddings
Layer 1: Captures basic syntax
Layer 2-10: Captures grammar, context
Layer 11-20: Captures meaning, semantics
Layer 21+: Prepares for output
# Final embeddings → Probability distribution over vocabulary
# Picks most likely next token, converts back to text
Output: "Whiskers soft as morning dew..."
Embedding Dimensions Explained
Most LLMs use 768, 1024, or larger dimensional embeddings. But what do these dimensions represent?
- Dimension 42 might respond to "plural" vs "singular"
- Dimension 103 might encode "positive" vs "negative" sentiment
- Dimension 256 might capture "past" vs "present" tense
- Dimension 512 might relate to "animate" vs "inanimate" objects
| Model | Embedding Size | Parameters |
|---|---|---|
| GPT-2 | 768 | 117M - 1.5B |
| BERT-base | 768 | 110M |
| GPT-3 | 12,288 | 175B |
| Claude 3 | ~4096-8192 (estimated) | Unknown |
Vector Embeddings in Prompt Engineering
Understanding embeddings helps you write better prompts!
1. Semantic Similarity in Context
LLMs understand your prompt by comparing it to patterns seen during training.
Example Prompts:
Prompt A: "Explain quantum physics to a 5-year-old"
Prompt B: "Describe quantum mechanics in simple terms for children"
Prompt C: "What is the weather today?"
Embedding Similarities:
| Comparison | Similarity | Why? |
|---|---|---|
| A ↔ B | 0.92 | Almost identical meaning! |
| A ↔ C | 0.15 | Completely different topics |
2. Few-Shot Learning Through Embeddings
When you provide examples in prompts, the LLM uses embeddings to find patterns.
Example Prompt:
Classify sentiment:
Text: "I love this product!" → Positive
Text: "This is terrible" → Negative
Text: "Amazing experience!" → Positive
Text: "Worst purchase ever" → Negative
Text: "Pretty good overall" → ?
What Happens Internally:
- Converts all texts to embeddings
- Finds "Pretty good overall" is closer to positive examples
- Outputs: "Positive"
3. Instruction Following
Instruction-tuned models learn to recognize instruction patterns through embeddings.
These all work similarly:
"Write a poem about..."
"Create a poem about..."
"Compose a poem about..."
"Generate a poem about..."
Why? Their embeddings are very similar! The model recognizes the pattern.
Retrieval-Augmented Generation (RAG) and Embeddings
RAG is one of the most important uses of embeddings in modern AI systems!
How RAG Works
Problem: LLM doesn't know about your company's internal documents.
Solution: Use embeddings to find relevant documents, then feed them to the LLM!
Document 1: "Our return policy is 30 days"
→ Embedding: [0.3, -0.5, 0.7, ...]
Document 2: "Shipping takes 3-5 business days"
→ Embedding: [0.1, 0.4, -0.2, ...]
Document 3: "Contact support at help@company.com"
→ Embedding: [-0.2, 0.6, 0.3, ...]
User: "What's your return policy?"
→ Embedding: [0.2, -0.4, 0.8, ...]
Compare user question embedding with all document embeddings:
- Doc 1 similarity: 0.89 ← MOST SIMILAR!
- Doc 2 similarity: 0.23
- Doc 3 similarity: 0.15
Retrieve Document 1
Context: "Our return policy is 30 days"
Question: "What's your return policy?"
Answer: Based on the context, our return policy is 30 days.
Real-World Example:
# Simplified RAG implementation concept
def answer_question(user_question):
# 1. Convert question to embedding
question_embedding = embed(user_question)
# 2. Find most similar documents
similarities = []
for doc in document_database:
similarity = cosine_similarity(question_embedding, doc.embedding)
similarities.append((doc, similarity))
# 3. Get top 3 most relevant
top_docs = sorted(similarities, key=lambda x: x[1], reverse=True)[:3]
# 4. Create prompt with context
context = "\n".join([doc.text for doc, _ in top_docs])
prompt = f"Context: {context}\n\nQuestion: {user_question}\n\nAnswer:"
# 5. Get LLM response
return llm.generate(prompt)
Semantic Search with Embeddings
Traditional search uses keyword matching. Semantic search uses embeddings to understand meaning!
Keyword Search (Old Way)
User searches: "best laptop for programming"
Traditional search finds: Documents containing exact words "best", "laptop", "programming"
Misses: "top computer for coding" (different words, same meaning!)
Semantic Search (Embedding Way)
User searches: "best laptop for programming"
Embedding search finds:
- "Best laptop for programming" (exact match) - 0.99 similarity
- "Top computer for coding" (same meaning!) - 0.87 similarity
- "Great machine for developers" (similar meaning) - 0.82 similarity
- "Powerful notebook for software development" - 0.78 similarity
Computing Embeddings - Practical Code
Let's see how to actually work with embeddings in practice:
Using OpenAI's Embedding API
import openai
import numpy as np
# Get embedding for text
def get_embedding(text):
response = openai.Embedding.create(
input=text,
model="text-embedding-ada-002" # 1536 dimensions
)
return response['data'][0]['embedding']
# Example usage
text1 = "The cat sat on the mat"
text2 = "A feline rested on the rug"
text3 = "Python is a programming language"
embedding1 = get_embedding(text1)
embedding2 = get_embedding(text2)
embedding3 = get_embedding(text3)
print(f"Embedding dimension: {len(embedding1)}")
print(f"First 5 values: {embedding1[:5]}")
Output:
Embedding dimension: 1536
First 5 values: [0.0023, -0.0091, 0.0047, -0.0234, 0.0156]
Computing Similarity
def cosine_similarity(vec1, vec2):
"""Calculate cosine similarity between two vectors"""
vec1 = np.array(vec1)
vec2 = np.array(vec2)
dot_product = np.dot(vec1, vec2)
norm1 = np.linalg.norm(vec1)
norm2 = np.linalg.norm(vec2)
return dot_product / (norm1 * norm2)
# Compare similarities
sim_1_2 = cosine_similarity(embedding1, embedding2)
sim_1_3 = cosine_similarity(embedding1, embedding3)
print(f"Similarity (cat sentence vs feline sentence): {sim_1_2:.3f}")
print(f"Similarity (cat sentence vs Python sentence): {sim_1_3:.3f}")
Output:
Similarity (cat sentence vs feline sentence): 0.892
Similarity (cat sentence vs Python sentence): 0.134
The cat sentences are 89% similar, but cat and Python are only 13% similar! 🎯
3. Embedding-Based Prompt Optimization
Find the optimal prompt by testing variations and measuring output quality:
prompt_variations = [
"Write a professional email about {topic}",
"Compose a formal email regarding {topic}",
"Draft a business email on {topic}",
"Create a professional message about {topic}"
]
# Test each and measure output quality
# (quality could be measured by human evaluation or another LLM)
# Find which prompt embedding leads to best results
# Use that pattern for future prompts
Common Mistakes and Best Practices
Mistakes to Avoid
# WRONG - embeddings from different models aren't comparable!
bert_embedding = get_bert_embedding("cat")
gpt_embedding = get_gpt_embedding("cat")
similarity = cosine_similarity(bert_embedding, gpt_embedding) # Meaningless!
# CORRECT - use same model
emb1 = get_gpt_embedding("cat")
emb2 = get_gpt_embedding("dog")
similarity = cosine_similarity(emb1, emb2) # Valid!
If you update your LLM, regenerate all embeddings! Different model versions produce different embeddings.
Most embedding models have a max token limit (e.g., 512 tokens). Don't exceed it!
# WRONG
very_long_text = "..." * 10000 # Way too long!
embedding = get_embedding(very_long_text) # Gets truncated!
# CORRECT
chunks = split_text_into_chunks(very_long_text, max_tokens=500)
embeddings = [get_embedding(chunk) for chunk in chunks]
For faster similarity search, normalize embeddings to unit length:
# Better for large-scale search
def normalize(vec):
return vec / np.linalg.norm(vec)
embedding = normalize(get_embedding(text))
Best Practices
Embedding computation is expensive. Cache them for reuse!
import pickle
# Save embeddings
embeddings = {'text1': emb1, 'text2': emb2}
with open('embeddings.pkl', 'wb') as f:
pickle.dump(embeddings, f)
# Load later
with open('embeddings.pkl', 'rb') as f:
embeddings = pickle.load(f)
Process multiple texts at once for efficiency:
# Instead of this (slow):
for text in texts:
embedding = get_embedding(text)
# Do this (fast):
embeddings = get_embeddings_batch(texts) # Process all at once
Test similarity on known pairs to ensure embeddings work correctly:
# Test pairs that should be similar
assert cosine_similarity(
get_embedding("cat"),
get_embedding("feline")
) > 0.7 # Should be high similarity
# Test pairs that should be different
assert cosine_similarity(
get_embedding("cat"),
get_embedding("mathematics")
) < 0.3 # Should be low similarity
| Use Case | Recommended Model | Why? |
|---|---|---|
| General semantic search | text-embedding-ada-002 | Good balance of quality and speed |
| Multilingual | multilingual-e5-large | Supports 100+ languages |
| Code search | code-search-ada-code-001 | Optimized for programming |
| Domain-specific | Fine-tuned model | Best performance on specialized data |
Vector Databases for Embeddings
For production systems with millions of embeddings, use specialized vector databases:
Popular Vector Databases
| Database | Best For | Key Feature |
|---|---|---|
| Pinecone | Production apps | Fully managed, scalable |
| Weaviate | Hybrid search | Combines vector + keyword search |
| Milvus | Large scale | Open source, billions of vectors |
| Chroma | Prototyping | Simple, lightweight |
| FAISS | Research | Facebook's library, very fast |
Embeddings for Different Data Types
Text Embeddings
What we've discussed so far - converts words, sentences, paragraphs to vectors.
Image Embeddings
Models like CLIP create embeddings for images:
Image of cat → [0.5, -0.2, 0.8, ...]
Text "a cat" → [0.4, -0.3, 0.7, ...] # Similar to image embedding!
Applications:
- Image search using text queries
- Finding similar images
- Image classification
Multimodal Embeddings
Modern models create unified embeddings for text, images, and more!
Text: "A red car" → [0.3, 0.5, -0.2, ...]
Image: [photo of red car] → [0.2, 0.6, -0.1, ...] # Close to text!
- Search images using text descriptions
- Compare videos with text queries
- Generate images from text (DALL-E, Midjourney)
- Cross-modal understanding (text + image + audio)
Measuring Embedding Quality
Intrinsic Evaluation
Test on standard benchmarks:
| Benchmark | Tests |
|---|---|
| Word Similarity | Correlation with human judgments (SimLex-999) |
| Analogy | "king - man + woman = ?" (should get "queen") |
| Clustering | Do similar concepts cluster together? |
Extrinsic Evaluation
Test on your actual task:
- Search accuracy: Do users find what they want?
- RAG performance: Does it retrieve relevant documents?
- Classification accuracy: Does clustering work well?
Fine-Tuning Embeddings
For specialized domains, fine-tune embedding models on your data:
When to Fine-Tune
- Domain-specific terminology: Medical, legal, technical jargon
- Poor out-of-box performance: Generic models struggle with your data
- Large dataset available: You have thousands of examples
Practical Applications Summary
| Application | How Embeddings Help | Example |
|---|---|---|
| Semantic Search | Find by meaning, not keywords | Search "laptop" finds "notebook computer" |
| RAG Systems | Retrieve relevant context | Chatbot finds relevant docs to answer |
| Recommendation | Find similar items | "Users who liked X also liked Y" |
| Clustering | Group similar content | Auto-categorize support tickets |
| Duplicate Detection | Find near-duplicates | Remove duplicate questions in FAQ |
| Classification | Assign categories | Sentiment analysis, topic labeling |
Key Takeaways 📝
- Embeddings = Meaning as Numbers: Convert text to vectors that capture semantic meaning
- Similar Meaning = Similar Vectors: "cat" and "dog" have closer embeddings than "cat" and "car"
- Cosine Similarity Measures Closeness: Use it to compare embeddings (range: -1 to 1)
- Contextual > Static: Modern LLMs create different embeddings based on context
- Dimensions Capture Features: Each of 768+ dimensions encodes different semantic aspects
- RAG Needs Embeddings: Retrieve relevant documents by comparing embeddings
- Semantic Search > Keyword Search: Understand intent, not just match words
- Cache for Performance: Computing embeddings is expensive, reuse them!
- Same Model Required: Only compare embeddings from the same model
- Vector DBs for Scale: Use specialized databases for millions of embeddings
Quick Reference 📋
Common Embedding Models:
# OpenAI
text-embedding-ada-002 (1536 dims)
# Sentence Transformers
all-MiniLM-L6-v2 (384 dims)
all-mpnet-base-v2 (768 dims)
# Multilingual
multilingual-e5-large (1024 dims)
# Code
code-search-ada-code-001 (1536 dims)
Similarity Thresholds (typical):
0.9 - 1.0 : Nearly identical
0.7 - 0.9 : Very similar
0.5 - 0.7 : Somewhat related
0.3 - 0.5 : Loosely related
0.0 - 0.3 : Different
Happy learning! 🎯✨
Comments
Post a Comment