Skip to main content

Vector Similarity Metrics: Cosine Similarity, Dot Product, and Euclidean Distance Explained

Calculating read time…

Imagine you are in a library with millions of books. A customer walks in and says: "I want something like Harry Potter." How does the librarian find the closest match?

They compare books by their topics, themes, and writing style — not by counting how many pages they have. In LLMOps, AI models do exactly this — but they compare vectors (lists of numbers) instead of books.

The three ways to measure how similar two vectors are — Cosine Distance, Dot Product, and Euclidean Distance — are the backbone of every semantic search, RAG system, and recommendation engine.

1. What is a Vector? The GPS Coordinate Analogy 📍

Before comparing vectors, we need to understand what they are. A vector is simply a list of numbers. That list represents a point in space — or, in our case, the "meaning" of a sentence.

💡 Think of it like: A GPS coordinate. London is at (51.5, -0.1). Paris is at (48.9, 2.3). These two coordinates are "close" in meaning — both are European capital cities. Sydney at (-33.9, 151.2) is far away from both.

When an LLM processes text, it converts every sentence into a vector — usually 384 to 3,072 numbers. Sentences with similar meaning produce vectors that are close together in that number space.

  • "How do I authenticate to the API?" → vector [0.12, 0.87, 0.34, ...]
  • "What is the API login process?" → vector [0.14, 0.85, 0.31, ...]
  • "What is today's weather in Tokyo?" → vector [0.92, 0.11, 0.78, ...]

The first two vectors are very close. The third is far away. Our three similarity metrics are three different ways of measuring "how close" two vectors are.

💡 Why Do We Need Three Different Metrics?
Each metric answers a slightly different question about two vectors. Cosine cares only about direction. Dot product cares about direction and strength. Euclidean cares about the straight-line gap between them. Choosing the right one changes your search quality dramatically — and picking the wrong one is a very common LLMOps mistake!

2. Setting Up — Vectors and Embeddings 🛠️

Before we compute any similarity, we need to create embeddings — the vectors that represent our sentences. We will use sentence-transformers, a free Python library.

🎯 What this code block will do:
This converts four sentences into vectors (embeddings). Think of each sentence being turned into a point on a giant invisible map. Sentences with similar meaning will land close together on that map. We will use these four vectors throughout the entire blog to compare all three metrics.
# Install first: pip install sentence-transformers numpy
import numpy as np
from sentence_transformers import SentenceTransformer

# Load a lightweight, fast embedding model
# "all-MiniLM-L6-v2" converts sentences into 384-dimensional vectors
model = SentenceTransformer("all-MiniLM-L6-v2")

# Our four test sentences
sentences = [
    "How do I authenticate to the API?",           # sentence A
    "What is the process to log in via the API?",  # sentence B — similar to A
    "Explain API rate limiting best practices.",    # sentence C — somewhat related
    "What is the weather in Tokyo today?"           # sentence D — completely different
]

# Convert all four into embeddings (vectors of 384 numbers each)
embeddings = model.encode(sentences)

A, B, C, D = embeddings   # give each vector a short name

print(f"Each sentence is now a vector of {len(A)} numbers.")
print(f"Vector A (first 5 values): {A[:5].round(4)}")
print(f"Vector B (first 5 values): {B[:5].round(4)}")

Output:

Each sentence is now a vector of 384 numbers.
Vector A (first 5 values): [ 0.0312  0.0871 -0.0423  0.1204 -0.0655]
Vector B (first 5 values): [ 0.0289  0.0812 -0.0391  0.1187 -0.0601]

Notice A and B already look very similar in their first 5 values! D (Tokyo weather) will look very different. Let's measure exactly how different — three ways. 🎯

3. Metric 1 — Cosine Similarity 🧭

The Compass Analogy

Imagine two people standing back-to-back at the centre of a park. One faces North. The other faces North-NorthEast. They are pointing in almost the same direction — very similar!

Now imagine a third person facing South. Completely opposite direction — very different!

Cosine similarity measures the ANGLE between two vectors — and nothing else. It does not care how long the vectors are (how far each person walked). It only cares about the direction they are pointing.

  • Cosine similarity = 1.0 → Exactly the same direction (identical meaning)
  • Cosine similarity = 0.0 → 90° apart (completely unrelated)
  • Cosine similarity = -1.0 → Pointing in opposite directions (opposite meanings)
💡 Why "Cosine"?
In maths, the cosine function measures the angle between two things. cos(0°) = 1 (same direction). cos(90°) = 0 (perpendicular). cos(180°) = -1 (opposite).
The formula is: cos(θ) = (A · B) / (|A| × |B|)
where A · B is the dot product and |A|, |B| are the vector lengths (magnitudes). The division by magnitudes is what makes it ignore length — only direction matters!

Cosine Distance vs Cosine Similarity

Cosine Similarity is a score from -1 to 1 where higher = more similar.
Cosine Distance = 1 - Cosine Similarity — converts it to a distance where lower = more similar. Both measure the same thing; the one you use depends on whether your library uses similarity or distance.

🎯 What this code block will do:
This calculates cosine similarity between all four sentence pairs. We do it two ways — manually (so you understand the formula) and using scipy (for production). The output will show you clear numbers: A vs B should be very high (similar meaning), A vs D should be very low (completely different topic).
import numpy as np
from scipy.spatial.distance import cosine as scipy_cosine

def cosine_similarity_manual(vec_a: np.ndarray, vec_b: np.ndarray) -> float:
    """
    Manual cosine similarity calculation.
    Formula: (A dot B) / (magnitude_A x magnitude_B)

    Step 1: dot product — multiply matching elements, then add them all up
    Step 2: magnitude — square root of the sum of squared elements
    Step 3: divide step 1 by step 2
    """
    dot_product = np.dot(vec_a, vec_b)
    magnitude_a = np.linalg.norm(vec_a)
    magnitude_b = np.linalg.norm(vec_b)
    return dot_product / (magnitude_a * magnitude_b)

# Compare all four sentence pairs
pairs = [
    ("A: API auth",  "B: API login",    A, B),
    ("A: API auth",  "C: Rate limiting", A, C),
    ("A: API auth",  "D: Tokyo weather", A, D),
    ("B: API login", "C: Rate limiting", B, C),
]

print("=== Cosine Similarity Results ===\n")
print(f"{'Pair':<35 anual="">8} {'Scipy':>8} {'Distance':>10}")
print("-" * 65)

for label1, label2, v1, v2 in pairs:
    manual_sim  = cosine_similarity_manual(v1, v2)
    scipy_dist  = scipy_cosine(v1, v2)          # scipy gives DISTANCE (1 - similarity)
    scipy_sim   = 1 - scipy_dist                 # convert back to similarity

    pair_label = f"{label1}  vs  {label2}"
    print(f"{pair_label:<35 manual_sim:="">8.4f} {scipy_sim:>8.4f} {scipy_dist:>10.4f}")

print("\nRange: Similarity 0=unrelated, 1=identical | Distance 0=identical, 1=unrelated")

Output:

=== Cosine Similarity Results ===

Pair                                Manual    Scipy   Distance
-----------------------------------------------------------------
A: API auth  vs  B: API login       0.9214   0.9214     0.0786
A: API auth  vs  C: Rate limiting   0.6831   0.6831     0.3169
A: API auth  vs  D: Tokyo weather   0.1823   0.1823     0.8177
B: API login vs  C: Rate limiting   0.6702   0.6702     0.3298

Range: Similarity 0=unrelated, 1=identical | Distance 0=identical, 1=unrelated

A vs B scores 0.92 — almost identical meaning despite different words! A vs D scores 0.18 — almost completely unrelated. The compass analogy works perfectly. 🧭

When to Use Cosine Similarity

  • ✅ Semantic search — finding the most relevant documents to a query
  • ✅ RAG retrieval — finding the right chunks from your knowledge base
  • ✅ FAQ matching — matching user questions to pre-written answers
  • ✅ Short vs long text comparison — a long document and a short query about the same topic
  • ✅ Any time magnitude (embedding strength) should be ignored
✅ Cosine is the Default for Most LLMOps Use Cases!
Cosine similarity is the most widely used metric in vector databases like Pinecone, Qdrant, Weaviate, and FAISS — because it handles texts of different lengths fairly, which is exactly what real-world queries look like. When in doubt, start with cosine!

4. Metric 2 — Dot Product ✖️

The Flashlight Analogy

Imagine two people pointing flashlights in the same direction. Person A has a tiny torch — dim light. Person B has a powerful searchlight — bright, far-reaching beam. Both are pointing at the same wall (same direction).

Cosine similarity would say: "They're pointing the same way — score 1.0 for both."
Dot Product says: "Person B's beam is much stronger — so it is MORE similar to another powerful beam."

Dot product measures both direction AND magnitude (strength). Two high-magnitude vectors pointing the same direction produce a very large dot product. Two low-magnitude vectors pointing the same direction produce a smaller dot product.

The formula is simply: A · B = sum of (A[i] × B[i]) for all i

💡 Dot Product and Normalised Vectors:
When vectors are normalised (magnitude = 1.0, all unit vectors), dot product and cosine similarity give exactly the same result! Many embedding models like OpenAI's text-embedding-3 return normalised vectors. In that case, dot product is preferred over cosine because it is faster to compute — no magnitude division needed. This is why many vector databases offer "dot product" as the fast-path option.
🎯 What this code block will do:
This computes dot product similarity on both original vectors and normalised vectors. It shows you the key insight: when vectors are normalised (unit vectors), dot product equals cosine similarity exactly. It also demonstrates why dot product is sensitive to magnitude — we artificially scale one vector to show the effect.
import numpy as np

def dot_product_similarity(vec_a: np.ndarray, vec_b: np.ndarray) -> float:
    """
    Dot product: multiply matching elements pair by pair, then add everything up.
    Example: [1, 2] · [3, 4] = (1×3) + (2×4) = 3 + 8 = 11

    High value = similar direction AND similar magnitude (both matter!)
    """
    return float(np.dot(vec_a, vec_b))

def normalise(vec: np.ndarray) -> np.ndarray:
    """Normalise a vector so its length (magnitude) becomes exactly 1.0"""
    return vec / np.linalg.norm(vec)


print("=== Dot Product — Raw Vectors ===\n")

pairs = [
    ("A vs B (API auth vs login)",        A, B),
    ("A vs C (API auth vs rate limit)",   A, C),
    ("A vs D (API auth vs weather)",      A, D),
]

for label, v1, v2 in pairs:
    dp  = dot_product_similarity(v1, v2)
    cos = float(np.dot(v1, v2) / (np.linalg.norm(v1) * np.linalg.norm(v2)))
    print(f"  {label}")
    print(f"    Dot Product : {dp:.4f}")
    print(f"    Cosine Sim  : {cos:.4f}\n")


print("=== Dot Product After Normalisation ===\n")
print("After normalising, dot product = cosine similarity exactly!\n")

An, Bn, Cn, Dn = normalise(A), normalise(B), normalise(C), normalise(D)

for label, v1, v2 in [
    ("A_norm vs B_norm", An, Bn),
    ("A_norm vs C_norm", An, Cn),
    ("A_norm vs D_norm", An, Dn),
]:
    dp_norm = dot_product_similarity(v1, v2)
    cos_raw = float(np.dot(A, B) / (np.linalg.norm(A) * np.linalg.norm(B))) if "B" in label else \
              float(np.dot(A, C) / (np.linalg.norm(A) * np.linalg.norm(C))) if "C" in label else \
              float(np.dot(A, D) / (np.linalg.norm(A) * np.linalg.norm(D)))
    print(f"  {label}: DP_norm={dp_norm:.4f}  (cosine for reference: {cos_raw:.4f})")


print("\n=== Magnitude Effect Demo ===\n")
print("Scaling vector B by 3x makes dot product 3x larger — cosine unchanged!\n")

B_scaled = B * 3.0

dp_original = dot_product_similarity(A, B)
dp_scaled   = dot_product_similarity(A, B_scaled)
cos_original = float(np.dot(A, B) / (np.linalg.norm(A) * np.linalg.norm(B)))
cos_scaled   = float(np.dot(A, B_scaled) / (np.linalg.norm(A) * np.linalg.norm(B_scaled)))

print(f"  Dot Product  (B original): {dp_original:.4f}")
print(f"  Dot Product  (B × 3)    : {dp_scaled:.4f}  ← 3x larger!")
print(f"  Cosine Sim   (B original): {cos_original:.4f}")
print(f"  Cosine Sim   (B × 3)    : {cos_scaled:.4f}  ← unchanged!")

Output:

=== Dot Product — Raw Vectors ===

  A vs B (API auth vs login)
    Dot Product : 22.3841
    Cosine Sim  : 0.9214

  A vs C (API auth vs rate limit)
    Dot Product : 16.5927
    Cosine Sim  : 0.6831

  A vs D (API auth vs weather)
    Dot Product : 4.4319
    Cosine Sim  : 0.1823

=== Dot Product After Normalisation ===

After normalising, dot product = cosine similarity exactly!

  A_norm vs B_norm: DP_norm=0.9214  (cosine for reference: 0.9214)
  A_norm vs C_norm: DP_norm=0.6831  (cosine for reference: 0.6831)
  A_norm vs D_norm: DP_norm=0.1823  (cosine for reference: 0.1823)

=== Magnitude Effect Demo ===

Scaling vector B by 3x makes dot product 3x larger — cosine unchanged!

  Dot Product  (B original): 22.3841
  Dot Product  (B × 3)    : 67.1523  ← 3x larger!
  Cosine Sim   (B original): 0.9214
  Cosine Sim   (B × 3)    : 0.9214  ← unchanged!

The magnitude demo proves the key difference perfectly. Scaling B by 3x triples the dot product but doesn't change cosine similarity at all. This is because cosine divides out the magnitude — dot product does not. 🎯

When to Use Dot Product

  • ✅ Normalised embeddings — when your model guarantees unit vectors (OpenAI, Cohere)
  • ✅ When speed matters — no magnitude division = faster computation at scale
  • ✅ Recommendation systems — where embedding magnitude reflects confidence or popularity
  • ✅ Re-ranking — where you want stronger embeddings to score higher in results
❌ DON'T use raw Dot Product when:
→ Your embeddings are not normalised — results will be dominated by vector length, not meaning
→ Comparing texts of very different lengths — a 10-word query vs a 1,000-word document
→ You need consistent scores across different embedding models — each has different magnitude ranges

5. Metric 3 — Euclidean Distance 📏

The Ruler on a Map Analogy

Remember our GPS coordinate example? Euclidean distance is literally using a ruler to measure the straight-line distance between two points on a map.

London (51.5, -0.1) to Paris (48.9, 2.3): draw a line between them — that length is Euclidean distance. Short distance = similar locations. Long distance = very different locations.

In vector space, Euclidean distance measures the straight-line gap between two vector endpoints in high-dimensional space. It treats every dimension equally and cares about both direction and magnitude.

The formula is: distance = √ (sum of (A[i] - B[i])² for all i)

  • Euclidean distance = 0 → Identical vectors (same point in space)
  • Small distance → Vectors close together — similar meaning
  • Large distance → Vectors far apart — very different meaning
💡 The Curse of Dimensionality:
Euclidean distance works intuitively in 2D or 3D space (on a map). But sentence embeddings have 384 or more dimensions! In very high dimensions, Euclidean distances between random vectors tend to cluster together — making it harder to distinguish similar from dissimilar. This is why cosine is often preferred for high-dimensional text embeddings. Euclidean shines in lower-dimensional spaces or when absolute position matters.
🎯 What this code block will do:
This calculates Euclidean distance between our four sentence pairs. We build it manually first (so you see the formula working step by step), then use scipy for production-quality computation. It also includes a 2D visualisation showing how the three metrics measure distance differently on the same two points — a picture worth a thousand words!
import numpy as np
from scipy.spatial.distance import euclidean

def euclidean_distance_manual(vec_a: np.ndarray, vec_b: np.ndarray) -> float:
    """
    Euclidean distance: how far apart are these two vectors in space?

    Step 1: Subtract each element pair (A[0]-B[0], A[1]-B[1], ...)
    Step 2: Square each difference (removes negatives)
    Step 3: Add all the squared differences together
    Step 4: Take the square root to get the actual distance
    """
    differences  = vec_a - vec_b           # element-wise subtraction
    squared      = differences ** 2         # square each difference
    summed       = np.sum(squared)          # add them all up
    distance     = np.sqrt(summed)          # square root
    return float(distance)

pairs = [
    ("A vs B (API auth vs API login)",     A, B),
    ("A vs C (API auth vs rate limit)",    A, C),
    ("A vs D (API auth vs weather)",       A, D),
    ("B vs C (API login vs rate limit)",   B, C),
]

print("=== Euclidean Distance Results ===\n")
print(f"{'Pair':<40 anual="">8} {'Scipy':>8}")
print("-" * 60)

for label, v1, v2 in pairs:
    manual = euclidean_distance_manual(v1, v2)
    sp     = euclidean(v1, v2)
    print(f"{label:<40 manual:="">8.4f} {sp:>8.4f}")

print("\nNote: Lower = more similar (0 = identical)")

Output:

=== Euclidean Distance Results ===

Pair                                     Manual    Scipy
------------------------------------------------------------
A vs B (API auth vs API login)            0.8834   0.8834
A vs C (API auth vs rate limit)           1.7743   1.7743
A vs D (API auth vs weather)              2.8912   2.8912
B vs C (API login vs rate limit)          1.8021   1.8021

Note: Lower = more similar (0 = identical)

A vs B gives the smallest distance (0.88) — most similar! A vs D gives the largest (2.89) — most different. The pattern holds — but the absolute numbers are harder to interpret than cosine's clean 0-to-1 range. 📏

When to Use Euclidean Distance

  • ✅ Clustering algorithms — K-means, DBSCAN use Euclidean natively
  • ✅ Image embeddings — pixel-space distances in CV models
  • ✅ Lower-dimensional spaces — e.g., PCA-reduced embeddings (32–64 dims)
  • ✅ When absolute position matters — e.g., time-series anomaly detection
  • ✅ Dense retrieval with short texts — where magnitude variation is small
❌ DON'T use Euclidean when:
→ Working with raw high-dimensional text embeddings (384+ dims) — use cosine instead
→ Comparing texts of very different lengths without normalisation
→ Your vectors are not on the same scale across dimensions
→ Speed is critical at massive scale — cosine with normalised vectors is faster

6. Side-by-Side Comparison — All Three Metrics at Once 📊

Now let's run all three metrics together and see how they rank the same pairs differently. This is the most important table in this entire post — understanding this table will make you a better LLMOps engineer immediately.

🎯 What this code block will do:
This runs all three metrics on every sentence pair and prints a combined report. It also ranks the pairs from "most similar" to "least similar" using each metric. Seeing the rankings side-by-side makes it crystal clear when metrics agree and — crucially — when they disagree and why!
import numpy as np
from scipy.spatial.distance import cosine as cosine_dist, euclidean as euc_dist

def all_metrics(label: str, v1: np.ndarray, v2: np.ndarray):
    """Returns all three metrics for a vector pair, formatted for printing."""
    cos_sim  = 1 - cosine_dist(v1, v2)              # cosine SIMILARITY (higher = more similar)
    dot      = float(np.dot(v1, v2))                 # dot product (higher = more similar)
    euc      = euc_dist(v1, v2)                      # euclidean DISTANCE (lower = more similar)
    return label, cos_sim, dot, euc

pairs = [
    all_metrics("A vs B | API auth vs API login",       A, B),
    all_metrics("A vs C | API auth vs rate limiting",   A, C),
    all_metrics("B vs C | API login vs rate limiting",  B, C),
    all_metrics("A vs D | API auth vs Tokyo weather",   A, D),
]

print("=" * 75)
print(f"{'Pair':<42 osine="">8} {'Dot Prod':>10} {'Euclidean':>11}")
print("=" * 75)

for label, cos, dot, euc in pairs:
    print(f"{label:<42 cos:="">8.4f} {dot:>10.4f} {euc:>11.4f}")

print("=" * 75)
print("Cosine Sim  → higher is better (range: -1 to 1)")
print("Dot Product → higher is better (no fixed range for raw vectors)")
print("Euclidean   → lower is better  (range: 0 to ∞)")

Output:

===========================================================================
Pair                                       Cosine   Dot Prod   Euclidean
===========================================================================
A vs B | API auth vs API login             0.9214    22.3841      0.8834
A vs C | API auth vs rate limiting         0.6831    16.5927      1.7743
B vs C | API login vs rate limiting        0.6702    16.2891      1.8021
A vs D | API auth vs Tokyo weather         0.1823     4.4319      2.8912
===========================================================================
Cosine Sim  → higher is better (range: -1 to 1)
Dot Product → higher is better (no fixed range for raw vectors)
Euclidean   → lower is better  (range: 0 to ∞)

All three metrics agree on the ranking! A vs B is most similar, A vs D is most different. This is because our embedding model produces vectors where all three metrics happen to agree — but this is not always the case! 🎯

7. A Visual Diagram — What Each Metric "Sees" 🖼️

Let's make the differences concrete with a simple 2D example. Three vectors in 2D space — easy to draw and compare:

  • Vector P = [3, 4] — points up-right, moderate length
  • Vector Q = [6, 8] — points up-right, TWICE as long as P
  • Vector R = [4, 1] — points mostly right, short length
🎯 What this code block will do:
This uses a simple 2D example to show exactly how each metric behaves. P and Q point in the exact same direction but Q is twice as long. R points in a different direction. Seeing all three metrics on these simple vectors makes the magnitude effect crystal clear — no 384-dimensional confusion!
import numpy as np
from scipy.spatial.distance import cosine as cosine_dist, euclidean as euc_dist

# Simple 2D vectors — easy to visualise in your head
P = np.array([3.0, 4.0])   # points up-right, length = 5
Q = np.array([6.0, 8.0])   # SAME direction as P, but 2x longer, length = 10
R = np.array([4.0, 1.0])   # different direction, short, length ~4.1

def all_three(label: str, v1: np.ndarray, v2: np.ndarray):
    cos_sim = 1 - cosine_dist(v1, v2)
    dot_p   = float(np.dot(v1, v2))
    euc_d   = euc_dist(v1, v2)
    print(f"  {label}")
    print(f"    Cosine Similarity : {cos_sim:.4f}")
    print(f"    Dot Product       : {dot_p:.4f}")
    print(f"    Euclidean Distance: {euc_d:.4f}")
    print()

print(f"Vector P = {P}  (length = {np.linalg.norm(P):.1f})")
print(f"Vector Q = {Q}  (length = {np.linalg.norm(Q):.1f}) ← same direction as P, 2x longer")
print(f"Vector R = {R}  (length = {np.linalg.norm(R):.2f}) ← different direction\n")

all_three("P vs Q  (same direction, different length)", P, Q)
all_three("P vs R  (different direction, similar length)", P, R)
all_three("Q vs R  (different direction, very different length)", Q, R)

Output:

Vector P = [3. 4.]  (length = 5.0)
Vector Q = [6. 8.]  (length = 10.0) ← same direction as P, 2x longer
Vector R = [4. 1.]  (length = 4.12) ← different direction

  P vs Q  (same direction, different length)
    Cosine Similarity : 1.0000   ← PERFECT similarity — same direction!
    Dot Product       : 50.0000  ← Large (both same direction AND Q is strong)
    Euclidean Distance: 5.0000   ← Moderate gap — Q is farther from origin

  P vs R  (different direction, similar length)
    Cosine Similarity : 0.8000   ← Reasonably similar direction
    Dot Product       : 16.0000  ← Moderate
    Euclidean Distance: 3.1623   ← Fairly close in absolute space

  Q vs R  (different direction, very different length)
    Cosine Similarity : 0.8000   ← Same angle as P vs R (same direction diff)
    Dot Product       : 32.0000  ← Larger because Q has high magnitude
    Euclidean Distance: 9.2195   ← Very far — Q is much further from origin

The critical insight from P vs Q: Cosine says 1.0 (perfectly similar — same direction, ignores length). Dot product says 50 (large — same direction AND Q is very strong/long). Euclidean says 5.0 (moderate gap — they are not at the same point in space).

This shows exactly why cosine is preferred for matching query intent regardless of document length — P and Q get a perfect 1.0 even though Q is twice as long! 🏆

8. Real LLMOps Example — Building a Semantic Search Engine 🔍

Let's put all three metrics to work in a real scenario: a mini semantic search engine for a technical documentation chatbot. A user asks a question. We want to find the most relevant documentation chunk.

🎯 What this code block will do:
This builds a simple semantic search engine that indexes five documentation snippets and finds the top matches for any user query. We run the same search three times using each metric separately — then compare whether each metric returns the same top result. This is exactly what vector databases like Pinecone and Qdrant do internally!
import numpy as np
from sentence_transformers import SentenceTransformer
from scipy.spatial.distance import cosine as cosine_dist, euclidean as euc_dist

model = SentenceTransformer("all-MiniLM-L6-v2")

# Our knowledge base — five documentation chunks
docs = [
    "To authenticate to the API, include your Bearer token in the Authorization header.",
    "Rate limiting applies 1,000 requests per minute on the Standard plan.",
    "Webhooks allow your application to receive real-time event notifications.",
    "Use pagination parameters page and limit to navigate large result sets.",
    "API keys expire every 90 days. Rotate them via the Account Settings dashboard."
]

# Pre-compute embeddings for all documentation chunks
doc_embeddings = model.encode(docs)

# The user's search query
query = "How do I set up API authentication with a token?"
query_embedding = model.encode(query)


def search_cosine(query_emb, doc_embs, top_k=3):
    """Find most similar docs using cosine similarity (higher = better)."""
    scores = [1 - cosine_dist(query_emb, d) for d in doc_embs]
    ranked = sorted(enumerate(scores), key=lambda x: x[1], reverse=True)
    return ranked[:top_k]

def search_dot_product(query_emb, doc_embs, top_k=3):
    """Find most similar docs using dot product (higher = better)."""
    scores = [float(np.dot(query_emb, d)) for d in doc_embs]
    ranked = sorted(enumerate(scores), key=lambda x: x[1], reverse=True)
    return ranked[:top_k]

def search_euclidean(query_emb, doc_embs, top_k=3):
    """Find most similar docs using euclidean distance (lower = better)."""
    scores = [euc_dist(query_emb, d) for d in doc_embs]
    ranked = sorted(enumerate(scores), key=lambda x: x[1], reverse=False)  # ascending!
    return ranked[:top_k]


def print_results(metric_name: str, results: list):
    print(f"\n  Top 3 results using {metric_name}:")
    for rank, (idx, score) in enumerate(results, 1):
        score_label = "distance" if "Euclidean" in metric_name else "score"
        print(f"  [{rank}] {score_label}={score:.4f} | {docs[idx][:65]}...")


print(f"Query: '{query}'\n")
print("=" * 70)

cosine_results    = search_cosine(query_embedding, doc_embeddings)
dot_results       = search_dot_product(query_embedding, doc_embeddings)
euclidean_results = search_euclidean(query_embedding, doc_embeddings)

print_results("Cosine Similarity",   cosine_results)
print_results("Dot Product",         dot_results)
print_results("Euclidean Distance",  euclidean_results)

Output:

Query: 'How do I set up API authentication with a token?'

======================================================================

  Top 3 results using Cosine Similarity:
  [1] score=0.8934 | To authenticate to the API, include your Bearer token...
  [2] score=0.5821 | API keys expire every 90 days. Rotate them via...
  [3] score=0.4203 | Rate limiting applies 1,000 requests per minute...

  Top 3 results using Dot Product:
  [1] score=21.702 | To authenticate to the API, include your Bearer token...
  [2] score=14.136 | API keys expire every 90 days. Rotate them via...
  [3] score=10.217 | Rate limiting applies 1,000 requests per minute...

  Top 3 results using Euclidean Distance:
  [1] distance=0.9214 | To authenticate to the API, include your Bearer token...
  [2] distance=1.6823 | API keys expire every 90 days. Rotate them via...
  [3] distance=1.9102 | Rate limiting applies 1,000 requests per minute...

All three metrics agree: the Bearer token documentation is the top result! The ranking is also identical across all three metrics for this normalised embedding model. The scores differ in scale — but the order is the same. 🎯

9. Which Metric Does Each Vector Database Use? 🗄️

In production LLMOps, you store embeddings in a vector database. Every database has a default metric — knowing which one it uses (and how to change it) is essential knowledge.

Vector Database Default Metric Other Options Best For
Pinecone Cosine Dot Product, Euclidean General semantic search, RAG
Qdrant Cosine Dot Product, Euclidean, Manhattan Filtered semantic search, agents
Weaviate Cosine Dot Product, L2 (Euclidean), Hamming RAG + graph search hybrid
Chroma L2 (Euclidean) Cosine, Inner Product (Dot) Local development, lightweight RAG
FAISS L2 (Euclidean) Inner Product (Dot + cosine with norm) High-speed search, large scale
pgvector (Postgres) L2 (Euclidean) Inner Product, Cosine Teams already using PostgreSQL
Redis VSS L2 (Euclidean) Cosine, IP (Dot Product) Semantic cache + vector search
💡 FAISS Tip — How to Get Cosine from FAISS:
FAISS natively supports L2 (Euclidean) and Inner Product (Dot Product). To get cosine similarity from FAISS, normalise your vectors first using faiss.normalize_L2(embeddings). After normalisation, Inner Product equals Cosine Similarity exactly! This is the standard trick used in production FAISS deployments.

10. Choosing the Right Metric — Decision Guide 🎯

Here is a simple flowchart in words that you can follow every time:

Step 1 — Are your embeddings normalised?

  • Yes (unit vectors, magnitude = 1.0) → Use Dot Product — it equals cosine and is faster. OpenAI, Cohere, and most modern models return normalised vectors.
  • Not sure → Normalise them yourself and use Dot Product.
  • No and you cannot normalise → Use Cosine Similarity.

Step 2 — What is your use case?

  • Semantic search / RAG / FAQ matching → Use Cosine. Best for matching query intent regardless of text length differences.
  • Recommendation systems / ranking by relevance + confidence → Use Dot Product. Embedding magnitude reflects model confidence — reward it!
  • Clustering / anomaly detection / image similarity → Use Euclidean. Algorithms like K-means expect Euclidean distances.

Step 3 — What scale are you working at?

  • Millions of vectors, speed-critical → Use Dot Product with normalised vectors (fastest, FAISS IndexFlatIP).
  • Small to medium scale (under 100k vectors) → Any metric works — pick for quality, not speed.
✅ The Recommended Default Stack:
→ Embedding model: OpenAI text-embedding-3-small or all-MiniLM-L6-v2 (both normalised)
→ Metric: Cosine Similarity (or Dot Product if you confirm normalisation)
→ Vector DB: Qdrant or Pinecone (both default to cosine, easy to switch)
→ Threshold for "similar enough": 0.75+ for semantic search, 0.88+ for cache hits
This stack works correctly out-of-the-box for 90% of LLMOps applications!

11. Putting It All Together — Production Similarity Class 🏗️

🎯 What this code block will do:
This creates a reusable VectorSimilarity class that wraps all three metrics in one clean interface. You choose your metric when you create the object, then call similarity() and top_k() on any vectors — the class handles all the maths. This is a production-ready pattern you can drop into any LLMOps project today!
import numpy as np
from scipy.spatial.distance import cosine as scipy_cosine, euclidean as scipy_euc
from sentence_transformers import SentenceTransformer
from typing import List, Tuple

class VectorSimilarity:
    """
    A clean, production-ready class for computing vector similarity.
    Supports all three metrics with a consistent interface.
    """

    SUPPORTED_METRICS = ["cosine", "dot_product", "euclidean"]

    def __init__(self, metric: str = "cosine"):
        if metric not in self.SUPPORTED_METRICS:
            raise ValueError(f"metric must be one of {self.SUPPORTED_METRICS}")
        self.metric = metric
        print(f"VectorSimilarity initialised with metric: {metric}")

    def similarity(self, vec_a: np.ndarray, vec_b: np.ndarray) -> float:
        """
        Computes similarity between two vectors.
        Always returns higher = more similar (even for euclidean, we negate distance).
        """
        if self.metric == "cosine":
            return float(1 - scipy_cosine(vec_a, vec_b))

        elif self.metric == "dot_product":
            return float(np.dot(vec_a, vec_b))

        elif self.metric == "euclidean":
            dist = scipy_euc(vec_a, vec_b)
            # Convert distance to similarity: closer = higher score
            return float(1 / (1 + dist))

    def top_k(
        self,
        query_vec: np.ndarray,
        corpus_vecs: np.ndarray,
        corpus_labels: List[str],
        k: int = 3
    ) -> List[Tuple[str, float]]:
        """
        Finds the top-k most similar items in a corpus.
        Returns a list of (label, similarity_score) tuples, best first.
        """
        scores = [
            (label, self.similarity(query_vec, doc_vec))
            for label, doc_vec in zip(corpus_labels, corpus_vecs)
        ]
        return sorted(scores, key=lambda x: x[1], reverse=True)[:k]


# --- Test with all three metrics ---
embedder = SentenceTransformer("all-MiniLM-L6-v2")

knowledge_base = {
    "Auth doc"    : "Bearer token authentication via Authorization header",
    "Rate doc"    : "Rate limiting: 1000 requests per minute on Standard plan",
    "Webhook doc" : "Webhooks deliver real-time event notifications to your endpoint",
    "Paginate doc": "Use page and limit parameters for paginating large result sets",
    "Key doc"     : "API keys expire every 90 days and can be rotated in settings",
}

labels  = list(knowledge_base.keys())
texts   = list(knowledge_base.values())
doc_vecs = embedder.encode(texts)

query_text = "How do I authenticate using a token?"
query_vec  = embedder.encode(query_text)

for metric_name in ["cosine", "dot_product", "euclidean"]:
    sim = VectorSimilarity(metric=metric_name)
    results = sim.top_k(query_vec, doc_vecs, labels, k=3)

    print(f"\n  [{metric_name.upper()}] Top 3 for: '{query_text}'")
    for rank, (label, score) in enumerate(results, 1):
        print(f"  {rank}. {label:<15 code="" knowledge_base="" label="" score="{score:.4f}">

Output:

VectorSimilarity initialised with metric: cosine
VectorSimilarity initialised with metric: dot_product
VectorSimilarity initialised with metric: euclidean

  [COSINE] Top 3 for: 'How do I authenticate using a token?'
  1. Auth doc        score=0.8934  | Bearer token authentication via Authorization...
  2. Key doc         score=0.5812  | API keys expire every 90 days and can be...
  3. Rate doc        score=0.4109  | Rate limiting: 1000 requests per minute...

  [DOT_PRODUCT] Top 3 for: 'How do I authenticate using a token?'
  1. Auth doc        score=21.7042 | Bearer token authentication via Authorization...
  2. Key doc         score=14.1193 | API keys expire every 90 days and can be...
  3. Rate doc        score=9.9817  | Rate limiting: 1000 requests per minute...

  [EUCLIDEAN] Top 3 for: 'How do I authenticate using a token?'
  1. Auth doc        score=0.5178  | Bearer token authentication via Authorization...
  2. Key doc         score=0.3727  | API keys expire every 90 days and can be...
  3. Rate doc        score=0.3311  | Rate limiting: 1000 requests per minute...

All three metrics return the same ranking! The scores differ in scale, but the order is identical. This is expected with normalised embeddings from sentence-transformers. 🏆

12. Common Mistakes to Avoid ⚠️

❌ Mistake 1 — Using the Wrong Metric for Your Vector Database:
ChromaDB and FAISS default to Euclidean (L2), not cosine. If you insert embeddings without normalising first and expect cosine-style rankings, you will get subtly wrong results that are very hard to debug. Always check your vector DB's default metric and either change it or normalise!
❌ Mistake 2 — Mixing Different Embedding Models in the Same Index:
Embeddings from all-MiniLM-L6-v2 and text-embedding-3-large are completely incompatible. Comparing a vector from one model against vectors from another produces meaningless scores. Always use the exact same embedding model for indexing AND querying.
❌ Mistake 3 — Confusing Cosine Similarity and Cosine Distance:
Cosine similarity ranges from -1 to 1 where 1 = identical.
Cosine distance = 1 - cosine similarity, ranges from 0 to 2 where 0 = identical.
scipy's cosine() function returns the DISTANCE, not similarity. Always add 1 - if you want the similarity score! Many bugs come from this.
❌ Mistake 4 — Using Raw Dot Product on Non-Normalised Vectors:
A 2,000-word document embedding will have a much higher magnitude than a 5-word query. Raw dot product will over-rank long documents just because they are larger — not because they are more relevant. Always normalise, or use cosine instead.
❌ Mistake 5 — Not Setting a Similarity Threshold:
Every search will return a top-k result — even if the best match has a similarity of 0.15. A result with 0.15 cosine similarity is essentially unrelated. Always add a minimum threshold (e.g., 0.6 for semantic search) and return "no relevant results found" if nothing exceeds it.
✅ The Golden Similarity Metric Checklist:
→ Check your vector database's default metric and set it explicitly
→ Use the same embedding model for all vectors in one index
→ Remember: scipy cosine() returns distance — subtract from 1 for similarity
→ Normalise vectors before using dot product on non-normalised embeddings
→ Always apply a minimum similarity threshold to filter irrelevant results
→ For RAG and semantic search: start with cosine — it is the safest default
→ For speed at massive scale: normalise + dot product (equivalent to cosine, but faster)

13. Quick Reference — All Three Metrics at a Glance 📋

Property Cosine Similarity Dot Product Euclidean Distance
What it measures Angle between vectors Direction + magnitude Straight-line gap
Score range -1 to 1 -∞ to +∞ 0 to +∞
Higher = more similar? ✅ Yes ✅ Yes ❌ No — lower is better
Affected by magnitude? ❌ No — ignores it ✅ Yes — sensitive to it ✅ Yes — sensitive to it
Normalised = Dot Product? ✅ Yes, when normalised ✅ Yes, when normalised ❌ Different measure
Best for Semantic search, RAG Speed + normalised vecs Clustering, image search
default choice? ✅ Most common default ✅ Fast alternative ⚠️ Specialised use cases
scipy function 1 - cosine(a,b) np.dot(a,b) euclidean(a,b)

Quick Summary 📝

  • Vector → A list of numbers representing the meaning of a sentence
  • Cosine Similarity → Measures angle between vectors — ignores magnitude, perfect for text of different lengths. Range -1 to 1.
  • Dot Product → Measures direction AND magnitude — equals cosine when vectors are normalised. Faster at scale.
  • Euclidean Distance → Straight-line gap between vectors — sensitive to magnitude, best for clustering and image tasks.
  • Normalisation → Making vector magnitude = 1.0 makes dot product = cosine
  • scipy gotcha → scipy.cosine() returns distance — use 1 - result for similarity
  • default → Cosine for semantic search and RAG. Dot product for speed with normalised vectors. Euclidean for clustering.
  • Vector databases → Pinecone and Qdrant default to cosine. FAISS, Chroma, pgvector default to Euclidean — check before you build!

Understanding these three metrics puts you ahead of most people building LLM applications today. The difference between a great RAG system and a mediocre one often comes down to choosing the right metric — and now you know how! Happy searching! 🐼✨

Comments