Imagine you are a teacher marking exam papers. You have 500 students. You cannot read every single answer yourself — it would take forever! So you hire a very smart assistant who reads each answer, compares it to the correct one, and gives it a score from 0 to 10. 📝
That is exactly what DeepEval does for your AI model. Instead of you manually checking every single AI response, DeepEval automatically tests them, scores them, and tells you which ones passed and which ones failed — in seconds!
1. What is DeepEval?
DeepEval is a free, open-source testing framework built specifically for evaluating LLM (Large Language Model) applications. Think of it as Pytest — but for AI.
When you build a chatbot, a RAG system, or an AI agent, you need to regularly check: "Is my AI still giving good answers?" DeepEval automates that check with 50+ ready-made metrics.
- Is the answer correct? → DeepEval checks it ✅
- Is the AI making things up (hallucinating)? → DeepEval catches it ✅
- Is the retrieved context actually relevant? → DeepEval scores it ✅
- Did the AI agent use the right tool in the right order? → DeepEval measures it ✅
LLMs do not fail loudly. They do not throw an error when they give a wrong answer — they confidently say the wrong thing with a smile! Without evaluation, you have no idea if your AI is getting better or worse over time. DeepEval is your safety net. 🥅
2. The Big Problem DeepEval Solves 🔥
Here is a scenario every LLMOps engineer faces:
You update your system prompt. Or you switch from GPT-4 to Claude. Or you add more documents to your RAG database. The AI looks better to you in quick testing. But did it actually improve? Or did it secretly get worse on some questions?
Without DeepEval, you are guessing. With DeepEval, you run a full evaluation suite — hundreds of test cases, scored automatically — and get a clear pass/fail report in minutes.
💡 Think of it like: A car manufacturer does crash tests before releasing a new car model. DeepEval is your crash test suite for AI updates. 🚗💥
→ Catch AI regressions before they reach your users
→ Compare two model versions side by side with data
→ Run evaluations as part of your CI/CD pipeline automatically
→ Test RAG pipelines, chatbots, and agents with purpose-built metrics
→ Generate synthetic test data when you don't have enough real examples
3. Key Concepts You Must Know First 📖
Before we write any code, let's understand five simple ideas:
Concept 1 — Test Case
A test case is one single experiment. It has three main parts:
- input → What you asked the AI. ("What is your return policy?")
- actual_output → What the AI actually said.
- expected_output → What the AI should have said. (The ideal answer.)
Concept 2 — Metric
A metric is the measuring stick. It takes a test case, analyses it, and gives it a score from 0 to 1. Score of 1 = perfect. Score of 0 = completely wrong.
Concept 3 — Threshold
The threshold is the minimum passing score. A threshold of 0.7 means: "If the metric scores below 0.7, this test case FAILS." You decide what threshold is acceptable for your use case.
Concept 4 — LLM-as-a-Judge
Many DeepEval metrics use a powerful AI (like GPT-4 or Claude) to judge the output of your AI. This is called LLM-as-a-Judge. One AI evaluates another AI. The judge AI has seen millions of examples and gives scores very close to what a human expert would give.
Concept 5 — Retrieval Context
If you are using a RAG system (where the AI looks up documents before answering), you also pass in the retrieval_context — the chunks of document the AI used. DeepEval then checks if those chunks were actually useful. 📄
RAG stands for Retrieval-Augmented Generation. Instead of the AI relying only on what it learned during training, it first searches a database of your documents for relevant information, then uses that to answer the question. Like a student using a textbook during an open-book exam! 📚
4. Installation and Setup ⚙️
These are the two commands to get DeepEval installed and ready on your computer. The first line installs DeepEval using pip (Python's package installer). The second line tells DeepEval which AI model to use as the judge — in this case, we are using OpenAI's GPT. You only need to run these once.
# Step 1: Install DeepEval
pip install deepeval
# Step 2: Set your OpenAI API key so DeepEval can use GPT as the judge
# (You can also use Claude, Gemini, or even a local Ollama model)
export OPENAI_API_KEY="your-openai-api-key-here"
# Optional: Log into Confident AI (DeepEval's cloud dashboard) for visual reports
deepeval login
That is it! DeepEval is now ready to evaluate your AI. 🎉
You can use Anthropic's Claude as the evaluator model too!
Just set:
export ANTHROPIC_API_KEY="your-anthropic-key"Then pass
model="claude-opus-4-5" when creating any metric.
DeepEval supports OpenAI, Anthropic, Gemini, Azure, and local Ollama models.
5. Your Very First DeepEval Test 🧪
Let's start with the simplest possible test. We have a customer support chatbot. A user asks about the return policy. We want to check if the AI's answer is correct.
This creates your first ever evaluation test case. Imagine it like a classroom quiz: we give the AI a question (input), we know the right answer (expected_output), and we check what the AI actually said (actual_output). DeepEval then scores how correct the actual answer is, using a metric called GEval. If the score is 0.5 or above, the test passes!
import pytest
from deepeval import assert_test
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
def test_customer_support_response():
# Step 1: Define the metric — "Is the answer correct?"
# GEval uses an LLM-as-a-judge to score based on your custom criteria
correctness_metric = GEval(
name="Correctness",
criteria="Determine if the actual output is correct based on the expected output.",
evaluation_params=[
LLMTestCaseParams.ACTUAL_OUTPUT,
LLMTestCaseParams.EXPECTED_OUTPUT
],
threshold=0.5 # score must be >= 0.5 to pass
)
# Step 2: Create the test case — one question and answer pair
test_case = LLMTestCase(
input="What is your return policy?",
# What your AI actually responded — in real usage, this comes from your app
actual_output="You can return any item within 30 days for a full refund.",
# The gold-standard perfect answer
expected_output="We offer a 30-day full refund at no extra cost."
)
# Step 3: Run the evaluation — this checks if the test passes
assert_test(test_case, [correctness_metric])
Run this test from your terminal with:
deepeval test run test_customer_support.py
Output:
Running 1 test(s) in test_customer_support.py...
test_customer_support_response
Metrics Summary
┌─────────────┬───────┬───────────┬────────┐
│ Metric │ Score │ Threshold │ Status │
├─────────────┼───────┼───────────┼────────┤
│ Correctness │ 0.82 │ 0.50 │ PASSED │
└─────────────┴───────┴───────────┴────────┘
1 passed in 3.24s ✅
Score of 0.82 — well above the 0.5 threshold. The test passes! 🏆 Your AI gave a correct-enough answer about the return policy.
6. The Most Important DeepEval Metrics 📊
DeepEval comes with 50+ metrics built in. Here are the ones you will use most often, grouped by use case:
For General LLM Applications
| Metric | What It Checks | Score Means |
|---|---|---|
| GEval | Custom criteria — you write the rule in plain English | 0 = fails your criteria, 1 = perfectly meets it |
| AnswerRelevancy | Is the answer actually relevant to the question asked? | 0 = completely off-topic, 1 = perfectly on-topic |
| HallucinationMetric | Is the AI making up facts not in the provided context? | 0 = no hallucination, 1 = completely made up |
| BiasMetric | Does the response contain unfair bias? | 0 = no bias detected, 1 = heavily biased |
| ToxicityMetric | Does the response contain harmful or toxic language? | 0 = clean, 1 = highly toxic |
For RAG Systems (Retrieval-Augmented Generation)
| Metric | What It Checks | Think of it as... |
|---|---|---|
| Faithfulness | Does the answer stick to the retrieved documents? | Is the answer faithful to the source material? |
| ContextualRelevancy | Are the retrieved chunks actually relevant to the question? | Did the search engine find the right pages? |
| ContextualPrecision | Are the most relevant chunks ranked first? | Is the best result at the top of search? |
| ContextualRecall | Did we retrieve all the chunks needed to answer fully? | Did we find ALL the relevant pages, not just some? |
For AI Agents (Trend)
| Metric | What It Checks |
|---|---|
| TaskCompletion | Did the agent actually complete what was asked? |
| ToolCorrectness | Did the agent call the right tools in the right order? |
| ArgumentCorrectness | Were the arguments passed to tools valid and correct? |
| PlanAdherence | Did the agent follow its own plan without going off track? |
7. Testing a RAG Pipeline — Step by Step 🔍
RAG systems are the most common LLM application in production today. They look up documents, then generate an answer based on what they find. Testing them needs special metrics — let's see how.
Imagine this scenario: You built a chatbot for an online store. It looks up your product database before answering customer questions. You need to make sure it finds the right information AND uses it faithfully.
This runs four RAG-specific metrics on one test case at the same time. Think of it like checking a student's essay from four angles: Did they answer the question? Did they stick to the textbook? Did the teacher give them the right textbook pages? Did those pages contain everything needed? All four checks happen in one shot!
from deepeval import evaluate
from deepeval.metrics import (
AnswerRelevancyMetric,
FaithfulnessMetric,
ContextualRelevancyMetric,
ContextualRecallMetric
)
from deepeval.test_case import LLMTestCase
# The chunks of documents your RAG system retrieved
# In a real app, these come from your vector database
retrieved_chunks = [
"Our wireless headphones model XR-500 have 30 hours of battery life.",
"The XR-500 supports Bluetooth 5.3 and has active noise cancellation.",
"Charging takes 2 hours via USB-C. The headphones weigh 250 grams."
]
# Create the test case with all RAG-related fields filled in
test_case = LLMTestCase(
# What the user asked
input="How long does the XR-500 battery last and how long does it take to charge?",
# What your AI actually answered (in real usage: call your app and get this)
actual_output=(
"The XR-500 offers 30 hours of battery life. "
"Charging takes 2 hours using the USB-C cable."
),
# The ideal perfect answer (your golden reference)
expected_output=(
"The XR-500 has 30 hours of battery. It charges in 2 hours via USB-C."
),
# The document chunks your RAG system fetched — critical for RAG metrics
retrieval_context=retrieved_chunks
)
# Define all four RAG evaluation metrics
answer_relevancy = AnswerRelevancyMetric(threshold=0.7)
faithfulness = FaithfulnessMetric(threshold=0.7)
context_relevancy = ContextualRelevancyMetric(threshold=0.7)
context_recall = ContextualRecallMetric(threshold=0.7)
# Run all metrics at once and get a full report
results = evaluate(
test_cases=[test_case],
metrics=[answer_relevancy, faithfulness, context_relevancy, context_recall]
)
Output:
Evaluating 1 test case(s) with 4 metric(s)...
Metrics Summary
┌──────────────────────────┬───────┬───────────┬────────┐
│ Metric │ Score │ Threshold │ Status │
├──────────────────────────┼───────┼───────────┼────────┤
│ Answer Relevancy │ 0.96 │ 0.70 │ PASSED │
│ Faithfulness │ 0.94 │ 0.70 │ PASSED │
│ Contextual Relevancy │ 0.88 │ 0.70 │ PASSED │
│ Contextual Recall │ 0.91 │ 0.70 │ PASSED │
└──────────────────────────┴───────┴───────────┴────────┘
Overall: 4/4 metrics passed ✅
All four metrics passed! The RAG pipeline is working correctly — it retrieved the right information and used it faithfully in the answer. 🎉
8. Catching Hallucinations — The Most Critical Test 🚨
Hallucination is when an AI confidently says something that is not true.
It is the number one problem in production LLM systems.
DeepEval's HallucinationMetric catches this automatically.
💡 Think of it like: A student who reads the wrong chapter of a textbook but answers questions as if they studied the right one. They sound confident, but what they say is not supported by the source material.
This deliberately tests an AI that is making something up. The context says the headphones cost $199 — but the AI invented a "$149 sale price." The HallucinationMetric will catch this fabrication and give a high hallucination score. A high score here is BAD — it means the AI is lying!
from deepeval.metrics import HallucinationMetric
from deepeval.test_case import LLMTestCase
# What the AI was given as its source of truth
factual_context = [
"The XR-500 headphones are priced at $199.",
"They are available in black and white colours only.",
"Free shipping is available on all orders over $50."
]
# The AI's response — notice it invented a "$149 sale price" that is NOT in the context!
test_case = LLMTestCase(
input="How much do the XR-500 headphones cost?",
actual_output=(
"The XR-500 headphones cost $199 normally, "
"but they are currently on sale for $149 this week!" # HALLUCINATION!
),
# For HallucinationMetric, we pass context (not expected_output)
context=factual_context
)
# Create the hallucination metric
# NOTE: Lower score = less hallucination = GOOD
# Higher score = more hallucination = BAD
hallucination_metric = HallucinationMetric(threshold=0.5)
hallucination_metric.measure(test_case)
print(f"Hallucination Score : {hallucination_metric.score:.2f}")
print(f"Test Passed : {hallucination_metric.is_successful()}")
print(f"Reason : {hallucination_metric.reason}")
Output:
Hallucination Score : 0.67
Test Passed : False
Reason : The actual output introduces a "$149 sale price" claim
that is not supported by any of the provided context.
This constitutes a hallucinated fabrication.
The metric caught the hallucination! Score of 0.67 is above the 0.5 threshold, so the test FAILS — as it should. Your monitoring system would now alert you to investigate this response. 🚨
For most metrics, higher score = better (you want 1.0 for correctness).
But for HallucinationMetric and BiasMetric and ToxicityMetric — a higher score is WORSE! These measure the amount of a problem, not the quality of the answer. A hallucination score of 0 means the AI is perfectly factual. Always check metric documentation to confirm score direction!
9. Building Your Evaluation Dataset 📂
Running one test case at a time is useful for debugging.
But in production, you need to evaluate hundreds or thousands
of test cases at once to get statistically reliable results.
DeepEval's EvaluationDataset handles this beautifully.
This creates a collection of five test cases (like a question bank) and runs all of them through two metrics at the same time. Think of it like a teacher grading 5 exam papers at once instead of one by one. DeepEval then shows you a summary report of how many passed and how many failed.
from deepeval.dataset import EvaluationDataset
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric, GEval, LLMTestCaseParams
from deepeval import evaluate
# Create multiple test cases — in real production this would come from a database
test_cases = [
LLMTestCase(
input="What is your return policy?",
actual_output="You can return items within 30 days for a full refund.",
expected_output="30-day full refund policy, no extra costs."
),
LLMTestCase(
input="Do you offer free shipping?",
actual_output="Yes, free shipping is available on orders over $50.",
expected_output="Free shipping on orders over $50."
),
LLMTestCase(
input="What colours does the XR-500 come in?",
actual_output="The XR-500 is available in black, white, and red.",
expected_output="The XR-500 comes in black and white only." # AI said RED — wrong!
),
LLMTestCase(
input="How long is the warranty?",
actual_output="All products come with a 2-year manufacturer warranty.",
expected_output="Products are covered by a 2-year warranty."
),
LLMTestCase(
input="Can I exchange instead of return?",
actual_output="Yes, exchanges are accepted within the 30-day return window.",
expected_output="Exchanges are allowed within the 30-day return period."
),
]
# Package them into a dataset
dataset = EvaluationDataset(test_cases=test_cases)
# Define the metrics to run against every test case
answer_relevancy = AnswerRelevancyMetric(threshold=0.7)
correctness = GEval(
name="Correctness",
criteria="Determine if the actual output is factually correct based on the expected output.",
evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
threshold=0.7
)
# Run all test cases through all metrics — one command does everything!
evaluate(
test_cases=dataset,
metrics=[answer_relevancy, correctness]
)
Output:
Evaluating 5 test case(s) with 2 metric(s)...
Results
┌─────────────────────────────────────────┬────────┬────────────┐
│ Test Case Input │ Corr. │ Ans.Relev │
├─────────────────────────────────────────┼────────┼────────────┤
│ What is your return policy? │ 0.88 ✅ │ 0.94 ✅ │
│ Do you offer free shipping? │ 0.91 ✅ │ 0.97 ✅ │
│ What colours does the XR-500 come in? │ 0.31 ❌ │ 0.85 ✅ │ ← FAILED!
│ How long is the warranty? │ 0.87 ✅ │ 0.93 ✅ │
│ Can I exchange instead of return? │ 0.84 ✅ │ 0.92 ✅ │
└─────────────────────────────────────────┴────────┴────────────┘
Overall: 4/5 correctness passed | 5/5 relevancy passed
DeepEval found the problem instantly! The colour question failed because the AI said "red" — a colour that does not exist for that product. The dataset evaluation caught it. 🎯
10. G-Eval — Custom Criteria in Plain English ✍️
The most powerful and flexible metric in DeepEval is G-Eval. It lets you define any evaluation criteria in plain, everyday English — and uses an LLM-as-a-judge to score against that criteria.
💡 Think of it like: You are the teacher and you write the rubric yourself. G-Eval is the teaching assistant who grades all papers according to your rubric — with the intelligence of a senior examiner.
This creates three completely custom evaluation metrics — each one defined in plain English. No maths, no programming logic. You just describe what "good" looks like in words, and G-Eval scores the AI's response against that description. This is incredibly powerful for business-specific requirements!
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
from deepeval import evaluate
# Custom Metric 1: Is the tone friendly for a customer support bot?
friendliness_metric = GEval(
name="Friendly Tone",
criteria=(
"Evaluate if the response has a warm, helpful, and professional tone. "
"It should not sound robotic, cold, or dismissive. "
"It should make the customer feel heard and valued."
),
evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT],
threshold=0.7
)
# Custom Metric 2: Is the response concise (not too long, not too short)?
conciseness_metric = GEval(
name="Conciseness",
criteria=(
"Evaluate if the response answers the question directly without unnecessary "
"filler words, excessive repetition, or irrelevant information. "
"A concise response gets to the point quickly while being complete."
),
evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT],
threshold=0.7
)
# Custom Metric 3: Does the response include a clear call-to-action?
call_to_action_metric = GEval(
name="Clear Next Step",
criteria=(
"Evaluate if the response provides the customer with a clear next step "
"or action to take — such as a link, a phone number, or specific instructions. "
"A response without any guidance on what to do next scores low."
),
evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT],
threshold=0.6
)
# Test case to evaluate
test_case = LLMTestCase(
input="My order has been stuck on 'processing' for 5 days. What should I do?",
actual_output=(
"I completely understand how frustrating that must be — let's get this sorted! "
"Orders typically process within 1-2 business days, so a 5-day delay "
"definitely needs looking into. "
"Please visit our Order Status page at orders.shop.com and click 'Contact Support' "
"with your order number — our team will respond within 2 hours. "
"I'm sorry for the inconvenience! "
)
)
# Run all three custom metrics at once
evaluate(
test_cases=[test_case],
metrics=[friendliness_metric, conciseness_metric, call_to_action_metric]
)
Output:
Metrics Summary
┌──────────────────┬───────┬───────────┬────────┐
│ Metric │ Score │ Threshold │ Status │
├──────────────────┼───────┼───────────┼────────┤
│ Friendly Tone │ 0.93 │ 0.70 │ PASSED │
│ Conciseness │ 0.81 │ 0.70 │ PASSED │
│ Clear Next Step │ 0.95 │ 0.60 │ PASSED │
└──────────────────┴───────┴───────────┴────────┘
3/3 metrics passed ✅
All three custom business requirements are checked automatically. No manual reading required — G-Eval handles it all! 🏆
11. Evaluating AI Agents — Big Deal 🤖
The biggest trend in LLMOps is AI agents. An agent does not just answer questions — it takes actions. It calls tools, searches the web, books flights, writes code, sends emails.
Testing an agent is harder than testing a chatbot. You cannot just check the final answer — you also need to check how the agent reached that answer. Did it use the right tools? In the right order? With correct arguments?
This uses DeepEval's
@observe decorator to automatically trace
every step your agent takes — which tools it called and what happened.
DeepEval then runs TaskCompletion and ToolCorrectness metrics against the trace.
It is like installing a black box flight recorder in your agent — every action is logged
and then reviewed by the evaluation system!
from deepeval.tracing import observe, update_current_span
from deepeval.dataset import Golden, EvaluationDataset
from deepeval.metrics import TaskCompletionMetric, ToolCorrectnessMetric
# Step 1: Decorate your tools with @observe so DeepEval can trace them
@observe(type="tool")
def search_product_database(product_name: str):
"""Simulates searching a product database"""
products = {
"XR-500": {"price": 199, "stock": 15, "colour": ["black", "white"]},
"XR-300": {"price": 129, "stock": 0, "colour": ["black"]}
}
return products.get(product_name, {"error": "Product not found"})
@observe(type="tool")
def check_shipping_cost(order_total: float):
"""Simulates checking shipping cost"""
return {"shipping": 0 if order_total >= 50 else 9.99}
# Step 2: Decorate your agent with @observe too
@observe(type="agent")
def shopping_assistant(user_request: str):
"""
A simple shopping agent that looks up product info
and calculates the final cost including shipping.
"""
# The agent decides to search for the product first
product_info = search_product_database("XR-500")
# Then calculates total with shipping
shipping_info = check_shipping_cost(product_info["price"])
total = product_info["price"] + shipping_info["shipping"]
return (
f"The XR-500 costs ${product_info['price']}. "
f"Shipping: ${shipping_info['shipping']}. "
f"Total: ${total}."
)
# Step 3: Define agent evaluation metrics
task_completion = TaskCompletionMetric(
threshold=0.7,
model="gpt-4o"
)
# Step 4: Create a dataset and run the agent through evaluation
dataset = EvaluationDataset(goldens=[
Golden(input="How much will XR-500 headphones cost me including shipping?")
])
for golden in dataset.evals_iterator(metrics=[task_completion]):
shopping_assistant(golden.input)
print("Agent evaluation complete! Check your Confident AI dashboard for results.")
DeepEval automatically traces every tool call the agent makes, then checks whether the agent successfully completed its task. No manual inspection needed! 🎯
→
TaskCompletionMetric → Did the agent finish what was asked?→
ToolCorrectnessMetric → Did it use the right tools?→
ArgumentCorrectnessMetric → Were tool arguments valid and sensible?→
PlanAdherenceMetric → Did it follow its own reasoning plan?→
PlanQualityMetric → Was the agent's plan a good one in the first place?
12. Generating Synthetic Test Data 🧬
A common problem: you want to evaluate your AI but you do not have enough test questions. Writing hundreds of test cases by hand is tedious. DeepEval can generate test cases for you automatically from your own documents!
💡 Think of it like: You give DeepEval your product manual. It reads it and generates 100 realistic questions a customer might ask — along with the ideal answers — all automatically. Your test dataset grows from 10 to 1,000 with one command! 🚀
This reads your documents, then automatically generates realistic question-and-answer pairs (called "goldens") that represent what real users might ask. These become your evaluation dataset. It is like having an AI write your exam questions and answer key for you — saving hours of manual work!
from deepeval.synthesizer import Synthesizer
from deepeval.dataset import EvaluationDataset
# Create a synthesizer — this uses an LLM to generate test cases
synthesizer = Synthesizer()
# Option A: Generate from raw text (paste your documentation directly)
product_docs = [
"""
XR-500 Wireless Headphones Product Manual:
The XR-500 offers 30 hours of battery life on a single charge.
Charging takes 2 hours via the included USB-C cable.
The headphones feature active noise cancellation (ANC) and transparency mode.
Bluetooth 5.3 ensures stable connectivity up to 10 metres range.
Available in black and white. Retail price: $199.
30-day return policy applies. 2-year manufacturer warranty included.
Free shipping on all orders over $50.
"""
]
# Generate 5 question-answer pairs from the documentation
goldens = synthesizer.generate_goldens_from_docs(
document_paths=product_docs, # can also pass file paths
max_goldens_per_document=5,
include_expected_output=True # also generate the ideal answer
)
# Show what was generated
print(f"Generated {len(goldens)} test cases!\n")
for i, golden in enumerate(goldens, 1):
print(f"Test {i}:")
print(f" Q: {golden.input}")
print(f" A: {golden.expected_output}")
print()
Output:
Generated 5 test cases!
Test 1:
Q: How long does the XR-500 battery last on a full charge?
A: The XR-500 provides 30 hours of battery life per charge.
Test 2:
Q: What cable is used to charge the XR-500?
A: The XR-500 charges via USB-C, taking 2 hours for a full charge.
Test 3:
Q: What colours are available for the XR-500?
A: The XR-500 is available in black and white.
Test 4:
Q: Does the XR-500 have noise cancellation?
A: Yes, the XR-500 features active noise cancellation (ANC) and transparency mode.
Test 5:
Q: What is the Bluetooth range on the XR-500?
A: The XR-500 uses Bluetooth 5.3 with a range of up to 10 metres.
From one paragraph of documentation, DeepEval generated five ready-to-use test cases! Scale this to your full documentation and you have hundreds of tests in minutes. 🎉
13. Integrating DeepEval into CI/CD — Automated Quality Gates 🚦
The real power of DeepEval in LLMOps comes when you run it automatically — every time you change your AI system. This is called putting DeepEval into your CI/CD pipeline.
CI/CD stands for Continuous Integration / Continuous Deployment. It means: every time you push code to GitHub, a robot automatically runs all your tests. If they pass, the code is deployed. If they fail, it is blocked.
DeepEval plugs directly into Pytest — so your AI evaluations run alongside your regular unit tests automatically! 🤖
This is a complete Pytest test file for your AI application. When you run
deepeval test run or just pytest,
it automatically evaluates all three test cases. If any metric score
falls below the threshold, the test fails — just like a unit test failing —
and your CI/CD pipeline blocks the deployment. Your users are protected!
# File: test_ai_quality.py
# Run with: deepeval test run test_ai_quality.py
import pytest
from deepeval import assert_test
from deepeval.metrics import (
GEval,
AnswerRelevancyMetric,
HallucinationMetric
)
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
# In a real app, you would import your actual LLM app function here
# from my_app import get_ai_response
# For this demo, we use hardcoded outputs
# Define reusable metrics — define once, use in many tests
CORRECTNESS = GEval(
name="Correctness",
criteria="Check if actual output is factually correct based on expected output.",
evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
threshold=0.7
)
RELEVANCY = AnswerRelevancyMetric(threshold=0.7)
HALLUCINATION = HallucinationMetric(threshold=0.4) # fail if hallucination > 40%
# Test 1: Return policy question
def test_return_policy():
tc = LLMTestCase(
input="What is the return policy?",
actual_output="Items can be returned within 30 days for a full refund.",
expected_output="30-day full refund with no extra costs.",
context=["All customers get a 30-day full refund at no extra cost."]
)
assert_test(tc, [CORRECTNESS, RELEVANCY, HALLUCINATION])
# Test 2: Product price question
def test_product_price():
tc = LLMTestCase(
input="How much does the XR-500 cost?",
actual_output="The XR-500 headphones are priced at $199.",
expected_output="The XR-500 retails for $199.",
context=["XR-500 headphones retail price is $199."]
)
assert_test(tc, [CORRECTNESS, RELEVANCY, HALLUCINATION])
# Test 3: Shipping question
def test_shipping_policy():
tc = LLMTestCase(
input="Is shipping free?",
actual_output="Free shipping is available on orders over $50.",
expected_output="Free shipping on orders above $50.",
context=["Free standard shipping is offered on all orders exceeding $50."]
)
assert_test(tc, [CORRECTNESS, RELEVANCY, HALLUCINATION])
Output when you run deepeval test run test_ai_quality.py:
Collecting tests...
test_return_policy ✅ PASSED (Correctness: 0.87 | Relevancy: 0.93 | Halluc: 0.05)
test_product_price ✅ PASSED (Correctness: 0.91 | Relevancy: 0.96 | Halluc: 0.02)
test_shipping_policy ✅ PASSED (Correctness: 0.89 | Relevancy: 0.95 | Halluc: 0.03)
3 passed in 8.47s ✅
All tests green! This is your AI quality gate. If you change your prompt and correctness drops to 0.4, the pipeline blocks the deployment automatically. 🛡️
14. DeepEval Dashboard — Visual Reports on Confident AI 📈
Running tests in the terminal is great for developers. But product managers, team leads, and stakeholders need something visual. DeepEval's cloud platform Confident AI gives you beautiful dashboards and comparison reports — for free!
What you get on the Confident AI dashboard:
- Test run history → See all past evaluation runs, compare scores over time, and spot regressions at a glance
- Side-by-side comparison → Compare Model A vs Model B test results with visual charts — perfect for A/B testing prompt changes
- Failed test details → Click any failed test to see the exact input, output, and reason the metric gave for the failure
- Team collaboration → Share reports with non-technical stakeholders who cannot read terminal output
These two lines connect your local DeepEval test runs to the Confident AI cloud dashboard. After running this, every evaluation you do locally will automatically upload its results to the cloud dashboard — no extra code needed!
# Step 1: Create a free account at confident-ai.com
# Then get your API key from the dashboard settings
# Step 2: Log in from your terminal — this connects your machine to the cloud
deepeval login --confident-api-key "your-confident-ai-key-here"
# Step 3: Now run your tests as normal — results auto-upload to the cloud!
deepeval test run test_ai_quality.py
# Your results will appear at:
# https://app.confident-ai.com/dashboard
# with full charts, comparisons, and failure analysis
→ Unlimited test runs stored in the cloud
→ Side-by-side model comparison
→ Shareable HTML reports
→ Team access (invite collaborators)
→ No credit card required to start
15. Common Mistakes to Avoid ⚠️
A response can be relevant but hallucinated. It can be faithful but too vague. Always run at least 2–3 complementary metrics together for a complete picture. Never judge AI quality by a single score.
A threshold of 0.3 means your AI passes even with terrible responses. Start at 0.7 for most metrics in production. Only lower the threshold if you have a strong, specific reason to do so.
Evaluating 5 questions and declaring your AI "production ready" is dangerous. For a real product, you need at least 50–200 test cases covering different question types, edge cases, and failure modes. Use the synthetic data generation to scale up your dataset quickly.
For HallucinationMetric, BiasMetric, and ToxicityMetric — a LOW score is GOOD and a HIGH score is BAD. This is opposite to metrics like AnswerRelevancy where high = good. Always check: is this metric measuring quality (high=good) or a problem (low=good)?
Every time you change your prompt, swap your LLM, update your RAG documents, or adjust your system configuration — re-run your full evaluation suite. What worked before may not work after a change. Evaluation must be continuous, not a one-time activity.
→ Run at least 2–3 metrics per test case
→ Set thresholds at 0.7+ for production quality gates
→ Have 50+ test cases before calling your AI production-ready
→ Include hallucination tests for every RAG pipeline
→ Add DeepEval to your CI/CD so evaluation runs on every code change
→ Use synthetic data generation to keep your test dataset growing
→ Use Confident AI dashboard to compare evaluation runs over time
16. DeepEval vs Other Tools — When to Use What 🛠️
| Tool | Best For | Free? | DeepEval Advantage |
|---|---|---|---|
| DeepEval ✅ | Full LLM evaluation suite, RAG, agents, CI/CD integration | ✅ Open source | 50+ metrics, Pytest integration, synthetic data gen, agent tracing |
| LangSmith | LangChain-native tracing and prompt A/B testing | ✅ Free tier | Best if you are already using LangChain |
| MLflow LLM Evaluate | Teams already using MLflow for experiment tracking | ✅ Open source | Best for teams that want one tool for ML + LLM |
| Evidently AI | Data drift + model monitoring in production | ✅ Open source | Best for ongoing monitoring, not one-off evaluation |
| Weights & Biases | Experiment tracking with visual dashboards | ✅ Free for individuals | Best for research teams comparing many model versions |
Start with DeepEval for evaluation + Confident AI for the free dashboard. Once you grow, add Evidently AI for production monitoring. All three are free to start and work beautifully together.
Great AI is not just about building a model — it is about knowing your model is great, with data to back it up. DeepEval gives you that confidence. Happy evaluating! 🐼✨
Comments
Post a Comment