Skip to main content

Langfuse in MLOps — Your AI Application's Black Box Recorder 🪢

Calculating read time…

Imagine you built a self-driving car but had absolutely no dashcam, no speed sensor, and no error log. If it crashes, how would you ever know why?

That's exactly the situation most developers face when they ship an LLM application without proper observability. The model gives a wrong answer — but you have no idea what prompt was sent, how long it took, how much it cost, or which step went wrong.


💡 Langfuse is your LLM application's dashcam, black box recorder, and co-pilot dashboard — all in one open-source tool.

What Exactly is Langfuse?

Langfuse is an open-source LLM Engineering Platform. It helps you observe, debug, evaluate, and improve AI applications built with Large Language Models (LLMs) like GPT-4, Claude, Gemini, or any open-source model.

Think of it like Google Analytics — but instead of tracking website visitors, it tracks every single interaction your AI application has with the LLM: what was asked, what was answered, how long it took, and how much it cost.

  • 🔍 Tracing → Records every LLM call from start to finish
  • 📊 Metrics → Shows cost, latency, token usage, and quality scores
  • 🧪 Evaluations → Judges whether your AI output is actually good
  • 📝 Prompt Management → Version-controls your prompts like Git for code
  • 🛝 Playground → Test and iterate prompts without touching your codebase
  • 📦 Datasets → Build test datasets from real production traces

The best part? Langfuse is 100% open-source, self-hostable, and has a generous free cloud tier. You can get started in minutes! ⚡

Why Do You Even Need Langfuse?

Traditional software is like a calculator — give it 2+2, it always returns 4. You can unit test it perfectly. LLMs are nothing like that!

LLMs are non-deterministic. The same prompt can produce different outputs. Complex applications chain multiple LLM calls together. A RAG chatbot might do 5 steps: retrieve documents, re-rank them, compress context, call the LLM, and format the output. If the final answer is wrong — which step failed?

💡 Without Langfuse: You're debugging blind. You literally guess.

💡 With Langfuse: You see a complete timeline of every step, every token, every cost — exactly like a flight data recorder after a crash.

Here are the most common problems Langfuse solves:

  • Hallucinations → The model made up facts. Which prompt caused it?
  • High cost → You're spending $500/month on tokens. Which feature costs the most?
  • Slow responses → Users complain about 8-second delays. Which step is the bottleneck?
  • Quality regression → A prompt change last Tuesday broke something. Which one?
  • Agent failures → Your AI agent took 12 steps and gave the wrong answer. What happened?

All of these are answered instantly by Langfuse traces. Let's see how it works! 🎯

Core Concept 1: Traces and Spans 🔍

Before writing any code, let's understand the two most important concepts in Langfuse. These mirror how hospitals track a patient's journey through the system.

What is a Trace?

A Trace represents one complete user interaction with your application — from the moment a user sends a message to the moment they receive the final response.

Think of it as one patient visit to a hospital. It has a start time, an end time, and it contains everything that happened in between.

What is a Span?

A Span is one individual step inside that trace — like the blood test, X-ray, or doctor consultation that happened during the patient visit.

A single trace can contain many nested spans. For a RAG chatbot, a trace might contain:

  • 📄 Span 1: User query received and processed
  • 🔎 Span 2: Vector database search (retrieval)
  • 🔢 Span 3: Document re-ranking
  • 🤖 Span 4: LLM call with context (the main generation)
  • ✅ Span 5: Response formatting and return

Langfuse records all of these spans, their exact timing, their inputs and outputs, and shows them in a beautiful tree view in the UI. 🌳

Setting Up Langfuse — Three Options 🛠️

Option A: Langfuse Cloud (Easiest — Recommended for Beginners)

Go to cloud.langfuse.com, sign up for free, and create a project. You'll get two keys: a Public Key and a Secret Key. That's it!

Option B: Self-Host with Docker (Free, Private, Full Control)

Run Langfuse entirely on your own machine with just two commands:

# Clone the repository
git clone https://github.com/langfuse/langfuse.git
cd langfuse

# Start Langfuse locally using Docker Compose
docker compose up

Open your browser at http://localhost:3000 and Langfuse is running locally! Your data never leaves your machine. Perfect for private or enterprise use.

Option C: Kubernetes / Helm (Production at Scale)

For large teams and high-traffic production systems, deploy via Helm chart on Kubernetes. This is the preferred method for enterprise-grade MLOps pipelines.

📌 Note: All code examples in this tutorial work with any of the three options above. You just change the LANGFUSE_BASE_URL environment variable to point to your cloud instance or self-hosted server.

Install the Python SDK

pip install langfuse

Set Your Environment Variables

import os

# Get these from your Langfuse project settings
os.environ["LANGFUSE_PUBLIC_KEY"] = "pk-lf-your-public-key"
os.environ["LANGFUSE_SECRET_KEY"] = "sk-lf-your-secret-key"

# Point to cloud or your self-hosted instance
os.environ["LANGFUSE_BASE_URL"]   = "https://cloud.langfuse.com"
# For self-hosted: "http://localhost:3000"

You're now connected to Langfuse! Let's send our first trace. 🎉

Your First Langfuse Trace — Hello World!

The simplest way to use Langfuse is the @observe() decorator. Just add it above any Python function and Langfuse automatically records everything that happens inside it — inputs, outputs, timing, and nested calls.

Example: Tracing a Simple OpenAI Call

from langfuse import observe
from langfuse.openai import openai   # Drop-in replacement for the standard openai import

# The @observe() decorator tells Langfuse to record this function as a trace
@observe()
def answer_user_question(user_question: str) -> str:
    """
    Answers a user question using GPT-4.
    Langfuse automatically captures: the prompt, the response,
    token usage, cost, and latency — all without any extra code!
    """

    response = openai.chat.completions.create(
        model    = "gpt-4o",
        messages = [
            {"role": "system", "content": "You are a helpful assistant. Be concise."},
            {"role": "user",   "content": user_question}
        ]
    )

    return response.choices[0].message.content


# Run it — Langfuse captures everything automatically!
answer = answer_user_question("What is machine learning in one sentence?")
print(answer)

Output:

Machine learning is a branch of artificial intelligence where systems
automatically learn patterns from data to make predictions or decisions
without being explicitly programmed for each task.

Now open your Langfuse dashboard and you'll see a trace appear with: the exact prompt sent, the exact response received, the model used, the number of tokens consumed, the cost in dollars, and the latency in milliseconds. 🎯

💡 Think of it like this: You wrote one line of code (@observe()) and instantly gained complete visibility into what your AI application is doing!

Tracing a Multi-Step RAG Pipeline 🔗

Real LLM applications aren't just one function — they're chains of steps. Let's trace a complete RAG (Retrieval-Augmented Generation) pipeline where multiple functions work together.

RAG is like a student answering an exam question. First they look up their notes (retrieval), then they read the relevant section (context building), then they write their answer (generation).

Example: Full RAG Pipeline with Nested Tracing

from langfuse import observe
from langfuse.openai import openai
from typing import List

# ── Step 1: Retrieve relevant documents ──────────────────────────────────────
@observe()
def retrieve_documents(query: str) -> List[str]:
    """
    Simulates a vector database search.
    In production, replace this with Qdrant, Pinecone, Weaviate, etc.
    Langfuse records: the query, results returned, and time taken.
    """

    # Simulated retrieved document chunks
    return [
        "MLOps is the practice of combining Machine Learning with DevOps principles.",
        "Key MLOps tools include MLflow, Kubeflow, and Langfuse for LLM applications.",
        "Langfuse provides tracing, evaluation, and prompt management for LLM pipelines."
    ]


# ── Step 2: Build the prompt with retrieved context ───────────────────────────
@observe()
def build_prompt(user_query: str, documents: List[str]) -> str:
    """
    Combines retrieved documents with the user question to create the final prompt.
    Langfuse records: the constructed prompt and its length.
    """

    context = "\n\n".join([f"- {doc}" for doc in documents])

    prompt = f"""
You are a helpful AI assistant. Use ONLY the context below to answer the question.
If the context doesn't contain the answer, say "I don't know."

Context:
{context}

Question: {user_query}

Answer:"""

    return prompt


# ── Step 3: Generate the answer using the LLM ────────────────────────────────
@observe()
def generate_answer(prompt: str) -> str:
    """
    Calls the LLM with the constructed prompt.
    Langfuse records: tokens used, cost, latency, and the model's response.
    """

    response = openai.chat.completions.create(
        model    = "gpt-4o",
        messages = [{"role": "user", "content": prompt}],
        max_tokens = 200
    )

    return response.choices[0].message.content.strip()


# ── Main pipeline: Langfuse records this as the parent trace ─────────────────
@observe()
def rag_pipeline(user_question: str) -> str:
    """
    The complete RAG pipeline.
    Langfuse shows this as one parent trace with three nested child spans:
    retrieve → build_prompt → generate_answer
    """

    print(f"\n Processing: '{user_question}'")

    # Each step is a nested span inside the parent trace
    docs   = retrieve_documents(user_question)
    prompt = build_prompt(user_question, docs)
    answer = generate_answer(prompt)

    print(f" Answer: {answer}\n")
    return answer


# ── Run it! ───────────────────────────────────────────────────────────────────
result = rag_pipeline("What is Langfuse used for in MLOps?")

Output:

 Processing: 'What is Langfuse used for in MLOps?'
 Answer: Langfuse is used in MLOps for LLM applications, specifically
providing tracing, evaluation, and prompt management capabilities.
It helps teams monitor and debug their LLM pipelines effectively.

In your Langfuse dashboard, you'll see one parent trace called rag_pipeline with three nested child spans underneath it — one for each function. You can click into any span to see exactly what went in and came out! 🌳

Adding Custom Metadata to Your Traces 🏷️

Raw traces are useful — but you become truly powerful when you attach custom metadata. This lets you filter, group, and analyze traces in ways that matter to your business.

Example: Enriching Traces with User and Session Info

from langfuse import observe, get_client

langfuse = get_client()

@observe()
def answer_with_metadata(
    user_question: str,
    user_id:       str,
    session_id:    str,
    product_area:  str
) -> str:
    """
    Same pipeline as before, but now we attach rich metadata.
    This lets you answer questions like:
    - Which user asked the most expensive questions?
    - Which product area has the highest latency?
    - Which session had the most failed responses?
    """

    # Update the current trace with metadata
    langfuse.update_current_trace(
        user_id  = user_id,                       # Track per-user cost and usage
        session_id = session_id,                  # Group multi-turn conversations
        tags     = [product_area, "rag", "v2"],   # Filter by feature/area
        metadata = {
            "product_area": product_area,
            "app_version":  "2.1.0",
            "environment":  "production"
        }
    )

    # The actual logic (same as before)
    docs   = retrieve_documents(user_question)
    prompt = build_prompt(user_question, docs)
    answer = generate_answer(prompt)

    return answer


# Run with metadata
result = answer_with_metadata(
    user_question = "Explain LLM observability in simple terms.",
    user_id       = "user-12345",
    session_id    = "session-abc-xyz",
    product_area  = "customer-support"
)
print(result)

What this gives you in the Langfuse dashboard:

  • 📊 Per-user cost analytics → See exactly how much user-12345 cost you this month
  • 💬 Session replay → See the entire multi-turn conversation of session-abc-xyz
  • 🏷️ Tag filtering → Filter all traces for "customer-support" to find issues in that feature
  • 🔖 Version tracking → See if app_version 2.1.0 performs better than 2.0.0

💡 Think of tags and metadata like labels on a filing cabinet. Without them, you'd have to open every drawer to find what you need. With them, you go straight to the right folder! 📁

Evaluations — Is Your AI Actually Any Good? 🧪

Tracing tells you what happened. Evaluations tell you whether it was good. These are two completely different things — and both are essential.

A trace might show that your LLM responded in 0.8 seconds using 350 tokens. That's great technically. But did it actually answer the user's question correctly? Was it factually accurate? Was it helpful? That's what evaluations measure.

Three Types of Evaluations in Langfuse:

  • 🤖 LLM-as-a-Judge → Use another LLM (like GPT-4) to score your output automatically. Fast and scalable.
  • 👤 Human Feedback → Collect thumbs-up/thumbs-down or star ratings from real users.
  • ✍️ Manual Labeling → Your team reviews specific traces and assigns quality scores.

Example: LLM-as-a-Judge Evaluation

from langfuse import get_client
import openai

langfuse = get_client()

def evaluate_answer_quality(
    trace_id:      str,
    user_question: str,
    ai_answer:     str
) -> dict:
    """
    Uses GPT-4 to evaluate the quality of an AI answer.
    Scores are sent back to Langfuse and attached to the original trace.

    This is called 'LLM-as-a-Judge' — one LLM evaluates another LLM's output.
    It's automated, scalable, and runs on every single production trace.
    """

    # Prompt the judge LLM to evaluate the answer
    judge_prompt = f"""
You are an expert evaluator for AI assistant responses.
Score the following answer on three dimensions, each from 0 to 10:

Question: {user_question}
Answer:   {ai_answer}

Respond with ONLY a JSON object like this (no extra text):
{{
  "relevance": <0-10>,
  "accuracy":  <0-10>,
  "clarity":   <0-10>,
  "reasoning": ""
}}
"""

    response = openai.chat.completions.create(
        model    = "gpt-4o",
        messages = [{"role": "user", "content": judge_prompt}]
    )

    import json
    scores = json.loads(response.choices[0].message.content)

    # Send scores to Langfuse — they appear in the trace and in analytics
    langfuse.score(
        trace_id = trace_id,
        name     = "relevance",
        value    = scores["relevance"] / 10,   # Normalize to 0.0 – 1.0
        comment  = scores["reasoning"]
    )
    langfuse.score(
        trace_id = trace_id,
        name     = "accuracy",
        value    = scores["accuracy"] / 10
    )
    langfuse.score(
        trace_id = trace_id,
        name     = "clarity",
        value    = scores["clarity"] / 10
    )

    print(f"✅ Evaluation complete for trace {trace_id}")
    print(f"   Relevance: {scores['relevance']}/10")
    print(f"   Accuracy:  {scores['accuracy']}/10")
    print(f"   Clarity:   {scores['clarity']}/10")
    print(f"   Reasoning: {scores['reasoning']}")

    return scores


# Example usage — pass the trace_id from a real Langfuse trace
evaluate_answer_quality(
    trace_id      = "trace-abc-12345",
    user_question = "What is Langfuse used for?",
    ai_answer     = "Langfuse is used for tracing and monitoring LLM applications."
)

Output:

✅ Evaluation complete for trace trace-abc-12345
   Relevance: 9/10
   Accuracy:  8/10
   Clarity:   9/10
   Reasoning: The answer is concise, accurate, and directly addresses the question
               though it could mention evaluation and prompt management as well.

Now in your Langfuse dashboard, every trace has quality scores attached. You can sort all your traces by "accuracy" and instantly find which ones scored lowest — those are the conversations your AI failed at! 🎯

Example: Collecting User Feedback (Thumbs Up/Down)

from langfuse import get_client

langfuse = get_client()

def record_user_feedback(trace_id: str, was_helpful: bool, comment: str = "") -> None:
    """
    Records a user's thumbs-up or thumbs-down on an AI response.
    This is the most valuable signal you can get — real human judgment.

    Call this whenever your UI's feedback button is clicked.
    The score appears in Langfuse alongside the original trace.
    """

    langfuse.score(
        trace_id = trace_id,
        name     = "user_feedback",
        value    = 1 if was_helpful else 0,   # 1 = helpful, 0 = not helpful
        comment  = comment or (" User marked as helpful" if was_helpful
                               else " User marked as unhelpful")
    )

    emoji = "👍" if was_helpful else "👎"
    print(f"{emoji} Feedback recorded for trace: {trace_id}")


# Simulating a user clicking thumbs-down and leaving a comment
record_user_feedback(
    trace_id   = "trace-abc-12345",
    was_helpful = False,
    comment    = "The answer was too vague and didn't mention the free tier."
)

Output:

👎 Feedback recorded for trace: trace-abc-12345

Over time, Langfuse builds a quality trend chart for your app. If user feedback scores drop after a deployment — you know a recent change broke something! 📉

Prompt Management — Version Control for Your Prompts 📝

Prompts are the most powerful lever in any LLM application. A small wording change can dramatically improve or tank your AI's performance. But where do most teams store prompts? Inside Python code. Inside .env files. In Notion docs.

This is a disaster. When something breaks, you can't trace which prompt version caused it. When the team collaborates, two people overwrite each other's work. There's no rollback, no history, no A/B testing.

💡 Langfuse Prompt Management fixes this. It gives you a centralized, version-controlled prompt registry that works like Git — for prompts!

Example: Creating and Versioning a Prompt

from langfuse import get_client

langfuse = get_client()

# ── Push a new prompt to Langfuse (do this via SDK or the UI) ────────────────
langfuse.create_prompt(
    name    = "customer-support-system",
    prompt  = """You are a friendly and professional customer support agent for TechCorp.

Your responsibilities:
- Answer questions about our products clearly and concisely.
- Always be polite, even if the customer is frustrated.
- If you don't know the answer, say: "Let me connect you with a specialist."
- Never make up pricing or feature information.

Respond in the same language the customer writes in.
Keep responses under 150 words unless more detail is truly necessary.""",
    labels  = ["production"],     # Mark this version as live in production
    config  = {
        "model":       "gpt-4o",
        "temperature": 0.3,
        "max_tokens":  300
    }
)

print(" Prompt version 1 pushed to Langfuse!")

Output:

 Prompt version 1 pushed to Langfuse!

Example: Fetching and Using a Prompt from Langfuse

from langfuse import get_client, observe
from langfuse.openai import openai

langfuse = get_client()

@observe()
def answer_support_question(customer_message: str, customer_name: str) -> str:
    """
    Fetches the live prompt from Langfuse (not from code!),
    compiles it with variables, and uses it to answer the customer.

    Why fetch from Langfuse instead of hardcoding?
    Because you can update the prompt in the UI and it goes live
    instantly — without redeploying your application! 🚀
    """

    # Fetch the production prompt from Langfuse
    prompt_template = langfuse.get_prompt("customer-support-system")

    # The prompt is automatically linked to this trace for analysis
    system_prompt = prompt_template.compile(customer_name=customer_name)

    response = openai.chat.completions.create(
        model    = prompt_template.config["model"],         # From Langfuse config
        messages = [
            {"role": "system", "content": system_prompt},
            {"role": "user",   "content": customer_message}
        ],
        temperature = prompt_template.config["temperature"],
        max_tokens  = prompt_template.config["max_tokens"]
    )

    return response.choices[0].message.content


# Run it!
reply = answer_support_question(
    customer_message = "Hi, I forgot my password and can't log in. Help!",
    customer_name    = "Priya"
)
print(reply)

Output:

Hi Priya! I'm sorry you're having trouble logging in.
To reset your password, please click "Forgot Password" on the login page
and follow the email instructions. If you don't receive the email within
5 minutes, check your spam folder or contact us at support@techcorp.com.
Happy to help with anything else!

Now in Langfuse, every trace is automatically linked to the exact prompt version that was used. If you update the prompt tomorrow, all future traces will link to version 2. You can compare performance between version 1 and version 2 directly in the dashboard! 📊

The Langfuse Dashboard — What You Actually See 📊

All the code above feeds data into the Langfuse UI. Let's tour what the dashboard actually shows you:

🗺️ Trace Explorer

A searchable, filterable list of every trace your application has generated. Click any trace to see its full nested span tree with timestamps, inputs, outputs, and scores. Share a trace URL with a teammate to debug together — no more screenshotting logs!

📈 Metrics Dashboard

Real-time charts showing: total traces over time, average latency, total token cost, quality score trends, and error rates. Slice everything by user, session, tag, or prompt version.

👤 User Analytics

See a breakdown per user: how many requests they made, total cost they generated, average quality of their responses, and a list of all their sessions. Instantly find your highest-cost users or your most frustrated users.

🧵 Session Replay

Group all traces from a multi-turn conversation into one session view. Watch the full back-and-forth between a user and your AI as it actually happened. This is invaluable for debugging complex chatbot failures.

🤖 Agent Graph View

For AI agents that call tools and make multiple decisions, Langfuse renders the entire execution as a graph diagram — showing which tools were called, in what order, and where it got stuck or failed.

📝 Prompt Manager

All your prompts in one place, with full version history. Edit prompts in the UI and promote a new version to production with one click. Roll back to the previous version if something goes wrong.

Integrations — Works With Everything! 🔌

Langfuse is framework-agnostic and integrates with the most popular tools in the LLM ecosystem. You don't need to switch your existing stack — just add Langfuse on top!

  • LLM Providers: OpenAI, Anthropic (Claude), Google Gemini, Mistral, Cohere, Ollama, any HuggingFace model, and 100+ others via LiteLLM.
  • LLM Frameworks: LangChain, LlamaIndex, Haystack, CrewAI, AutoGen, LangGraph.
  • Observability Standards: OpenTelemetry (OTel) — Langfuse is a native OTel receiver, so any OTel-compatible tool can send traces to it.
  • MLOps Tools: MLflow, Weights & Biases, Prefect, Airflow (via custom logging hooks).

Example: One-Line LangChain Integration

from langfuse.callback import CallbackHandler
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

# Create the Langfuse callback handler
langfuse_handler = CallbackHandler()

# Build a LangChain chain as usual
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant. Answer briefly."),
    ("user",   "{question}")
])

chain = prompt | ChatOpenAI(model="gpt-4o") | StrOutputParser()

# Pass the handler as a callback — Langfuse traces everything automatically!
result = chain.invoke(
    {"question": "What is LangChain used for?"},
    config={"callbacks": [langfuse_handler]}
)

print(result)
# Every step of the LangChain pipeline is now visible in Langfuse 🎉

Output:

LangChain is a framework for building applications powered by large language
models, enabling chaining of prompts, tools, memory, and data sources
into structured workflows.

Literally one line: config={"callbacks": [langfuse_handler]}. Your entire LangChain application is now fully observable! ⚡

Building Datasets for Offline Testing 📦

One of Langfuse's most powerful hidden features is the ability to build test datasets directly from your production traces.

Imagine you find 50 traces where users gave thumbs-down feedback. In Langfuse, you can click one button to add all those traces to a dataset. Now you have a real-world failure dataset — built from actual user interactions — that you can run tests against whenever you change your prompt or model.

Example: Creating a Dataset and Adding Items

from langfuse import get_client

langfuse = get_client()

# ── Create a named dataset ────────────────────────────────────────────────────
langfuse.create_dataset(
    name        = "support-failures-june-2025",
    description = "Real user interactions that received thumbs-down feedback in June 2025."
)

# ── Add test cases (input + expected output pairs) ────────────────────────────
test_cases = [
    {
        "input":    {"question": "How do I cancel my subscription?"},
        "expected": "Should explain the cancellation process clearly with specific steps."
    },
    {
        "input":    {"question": "Why was I charged twice?"},
        "expected": "Should acknowledge the billing issue and provide next steps."
    },
    {
        "input":    {"question": "Where is my order? It's been 3 weeks!"},
        "expected": "Should empathize with the delay and offer to escalate."
    }
]

for case in test_cases:
    langfuse.create_dataset_item(
        dataset_name     = "support-failures-june-2025",
        input            = case["input"],
        expected_output  = case["expected"]
    )

print(f"✅ Dataset created with {len(test_cases)} test cases.")

Output:

✅ Dataset created with 3 test cases.

Example: Running a Prompt Experiment on the Dataset

from langfuse import get_client, observe
from langfuse.openai import openai

langfuse = get_client()

@observe()
def run_with_new_prompt(question: str) -> str:
    """
    Runs our IMPROVED prompt (v2) against the failure dataset.
    We're testing whether v2 is better than v1 on the cases that previously failed.
    """
    response = openai.chat.completions.create(
        model    = "gpt-4o",
        messages = [
            {
                "role": "system",
                "content": """You are a compassionate, expert customer support specialist.
Always empathize with the customer first. Then provide clear, specific action steps.
If escalation is needed, proactively offer it. Be warm but professional."""
            },
            {"role": "user", "content": question}
        ]
    )
    return response.choices[0].message.content


# ── Run the new prompt against every dataset item ──────────────────────────────
dataset = langfuse.get_dataset("support-failures-june-2025")

for item in dataset.items:
    question = item.input["question"]

    # Run and link to the dataset item (so Langfuse knows this is a test run)
    with item.observe(run_name="v2-improved-prompt") as trace_id:
        answer = run_with_new_prompt(question)
        print(f"Q: {question}")
        print(f"A: {answer[:100]}...\n")

print("✅ Experiment complete! Check the Langfuse UI for side-by-side comparison.")

Output:

Q: How do I cancel my subscription?
A: I completely understand wanting to cancel, and I'm here to make this as smooth as possible.
   To cancel: Go to Settings → Billing → Cancel Subscription...

Q: Why was I charged twice?
A: I'm really sorry to hear about the double charge — that's frustrating and I want
   to resolve this right away. Please share your order ID and I'll...

Q: Where is my order? It's been 3 weeks!
A: Three weeks is absolutely too long, and I completely understand your frustration.
   Let me escalate this to our fulfillment team right now...

✅ Experiment complete! Check the Langfuse UI for side-by-side comparison.

In Langfuse, you can now see experiment results side-by-side: v1 vs v2. Which prompt got higher quality scores? Which was faster? Which was cheaper? This is how professional teams iterate on LLM applications systematically! 📊

Cost Monitoring — Stop Burning Money! 💸

Token costs can sneak up on you fast. A RAG application that sends 2,000 tokens per request, running 100,000 times a day, can cost thousands of dollars a month. Langfuse tracks every token, calculates the dollar cost, and lets you drill down to understand exactly where your budget is going.

What Langfuse Shows You in the Cost Dashboard:

  • 📅 Daily/weekly/monthly spend → Trending up? Something changed!
  • 👤 Cost per user → Which users generate the most tokens?
  • 🏷️ Cost per feature/tag → Which part of your app is most expensive?
  • 📝 Cost per prompt version → Did prompt v2 use more tokens than v1?
  • 🤖 Cost per model → Is GPT-4o 10x more expensive than GPT-4o-mini for this use case?

Example: Logging Cost Manually (for Custom/Local Models)

from langfuse import get_client

langfuse = get_client()

def log_custom_model_call(
    trace_id:       str,
    model_name:     str,
    prompt_tokens:  int,
    output_tokens:  int,
    cost_per_1k_prompt_tokens:  float,
    cost_per_1k_output_tokens:  float
) -> float:
    """
    Logs token usage and cost for a custom or self-hosted model.
    Useful for Ollama, vLLM, or any model not auto-tracked by Langfuse.
    """

    total_cost = (
        (prompt_tokens  / 1000) * cost_per_1k_prompt_tokens +
        (output_tokens  / 1000) * cost_per_1k_output_tokens
    )

    langfuse.score(
        trace_id = trace_id,
        name     = "custom_cost_usd",
        value    = total_cost,
        comment  = f"Model: {model_name} | Prompt: {prompt_tokens} tokens | Output: {output_tokens} tokens"
    )

    print(f"💰 Cost logged: ${total_cost:.6f} for {model_name}")
    return total_cost


# Example: Logging a Mistral-7B call running on local vLLM
log_custom_model_call(
    trace_id      = "trace-xyz-99999",
    model_name    = "mistral-7b-instruct",
    prompt_tokens  = 850,
    output_tokens  = 200,
    cost_per_1k_prompt_tokens  = 0.0002,   # Your actual compute cost
    cost_per_1k_output_tokens  = 0.0002
)

Output:

💰 Cost logged: $0.000210 for mistral-7b-instruct

Langfuse vs. Alternatives — How Does It Compare? 🆚

Langfuse isn't the only LLM observability tool out there. Here's how it compares:

  • Langfuse vs LangSmith: LangSmith is tightly integrated with LangChain (same team). Langfuse is framework-agnostic and works equally well with any LLM or framework. Langfuse is fully open-source; LangSmith has a proprietary core.
  • Langfuse vs Arize Phoenix: Phoenix is notebook-first and great for local debugging. Langfuse is production-first with a full cloud platform, team features, and prompt management.
  • Langfuse vs Helicone: Helicone works as a proxy (you redirect API calls through it). Langfuse uses SDKs and decorators — more flexible for complex multi-step applications.
  • Langfuse vs Datadog LLM Observability: Datadog is enterprise-grade and expensive. Langfuse is open-source and free to self-host. Both work well, but Langfuse is far more accessible for startups and individual developers.

💡 Bottom line: If you want open-source, self-hostable, full-featured, framework-agnostic LLM observability — Langfuse is currently the leading choice in 2025.

Best Practices ✅ and Common Mistakes ❌

Always do these things:

  • ✅ Add @observe() from day one → Don't wait until you have a problem. Set it up before your first deployment.
  • ✅ Add user_id to every trace → You'll thank yourself when you need to debug a specific user's issue.
  • ✅ Tag traces by feature → Use tags like "chat", "search", "summarizer" so you can filter by feature area.
  • ✅ Run evaluations on every production trace → Automated LLM-as-a-Judge catches quality drops before users notice.
  • ✅ Store prompts in Langfuse, not in code → This decouples prompt iteration from code deployment. A game-changer!
  • ✅ Build datasets from failed traces → Real-world failures make the best regression test cases.

Never do these things:

  • ❌ Don't skip tracing in development → You'll ship bugs to production you could have caught locally.
  • ❌ Don't log sensitive user data in traces → Always mask PII (names, emails, IDs) before it reaches Langfuse.
  • ❌ Don't ignore evaluation scores → Collecting scores but never acting on them is a waste. Review them weekly.
  • ❌ Don't use one giant function → Break your pipeline into small functions so each one gets its own span and is debuggable independently.
  • ❌ Don't forget to flush in batch jobs → In scripts that exit quickly, call langfuse.flush() at the end or traces may not all be sent.

Complete End-to-End Example — Production-Ready! 🏗️

Let's put everything together in one complete, production-ready script that covers tracing, metadata, evaluation, and cost tracking:

import os
import json
from langfuse import get_client, observe
from langfuse.openai import openai

langfuse = get_client()


@observe()
def retrieve_context(query: str) -> list:
    """Span 1: Retrieve relevant documents."""
    # Replace with your actual vector DB query (Qdrant, Pinecone, etc.)
    return [
        "Langfuse provides open-source LLM observability.",
        "It supports tracing, evals, prompt management, and datasets.",
        "Self-hostable via Docker in under 5 minutes."
    ]


@observe()
def build_rag_prompt(query: str, context_docs: list) -> str:
    """Span 2: Build the prompt with retrieved context."""
    context = "\n".join([f"• {doc}" for doc in context_docs])
    return f"Context:\n{context}\n\nQuestion: {query}\nAnswer:"


@observe()
def call_llm(prompt: str) -> str:
    """Span 3: Call the LLM."""
    response = openai.chat.completions.create(
        model    = "gpt-4o",
        messages = [{"role": "user", "content": prompt}],
        max_tokens = 200
    )
    return response.choices[0].message.content.strip()


@observe()
def auto_evaluate(trace_id: str, question: str, answer: str) -> None:
    """Span 4: Run LLM-as-a-Judge evaluation and send scores to Langfuse."""
    judge_prompt = f"""Rate this answer 1-10 for helpfulness. Reply with JSON only:
{{"score": <1-10>, "reason": ""}}

Question: {question}
Answer: {answer}"""

    result = openai.chat.completions.create(
        model    = "gpt-4o-mini",     # Use a cheaper model for evaluation
        messages = [{"role": "user", "content": judge_prompt}]
    )

    scores = json.loads(result.choices[0].message.content)
    langfuse.score(trace_id=trace_id, name="helpfulness", value=scores["score"] / 10)
    print(f"  📊 Helpfulness score: {scores['score']}/10 — {scores['reason']}")


# ── MAIN PIPELINE ─────────────────────────────────────────────────────────────

@observe()
def production_rag_pipeline(
    user_question: str,
    user_id:       str,
    session_id:    str
) -> str:
    """
    Complete production pipeline with:
    ✅ Multi-span tracing
    ✅ User and session metadata
    ✅ Automatic LLM-as-a-Judge evaluation
    ✅ Cost tracking (via Langfuse OpenAI integration)
    """

    # Attach metadata to the parent trace
    langfuse.update_current_trace(
        user_id    = user_id,
        session_id = session_id,
        tags       = ["rag", "production", "v3"]
    )

    # Execute the pipeline — each step is a nested span
    context_docs = retrieve_context(user_question)
    prompt       = build_rag_prompt(user_question, context_docs)
    answer       = call_llm(prompt)

    # Get the current trace ID for scoring
    trace_id = langfuse.get_current_trace_id()

    # Run evaluation asynchronously (or sync here for simplicity)
    auto_evaluate(trace_id, user_question, answer)

    print(f"\n✅ Answer: {answer}")
    return answer


# ── Run it! ───────────────────────────────────────────────────────────────────
result = production_rag_pipeline(
    user_question = "What makes Langfuse different from other observability tools?",
    user_id       = "user-98765",
    session_id    = "session-demo-001"
)

# IMPORTANT: In scripts, always flush to ensure all traces are sent before exit
langfuse.flush()

Output:

   📊 Helpfulness score: 9/10 — Clear, specific answer that directly addresses
     the key differentiators of Langfuse.

✅ Answer: Langfuse stands out because it's fully open-source and self-hostable,
   while offering enterprise-grade features like prompt management,
   LLM-as-a-Judge evaluations, and dataset management — all in one platform.
   Its framework-agnostic design means it works with any LLM or tool.

Open your Langfuse dashboard and you'll see the full trace with 4 nested spans, a helpfulness score of 0.9, user and session metadata, token costs, and latency breakdown. All from one decorator and a few lines of metadata! 🎉

Quick Summary 📝

What we learned today:

  • What Langfuse is → Open-source LLM engineering platform for observability, evals, and prompt management
  • Why it matters → Without it, debugging LLM apps is pure guesswork
  • Traces and Spans → Traces = one user request; Spans = individual steps inside it
  • Setup → Cloud (free), self-hosted Docker, or Kubernetes — all work the same way
  • @observe() decorator → One line that makes any Python function fully traceable
  • Evaluations → LLM-as-a-Judge, user feedback, and manual labeling
  • Prompt Management → Version-controlled prompts that update live without redeployment
  • Datasets → Build regression test datasets from real production failures
  • Cost Monitoring → Track every token and dollar by user, feature, and prompt version
  • Integrations → Works with OpenAI, Anthropic, LangChain, LlamaIndex, and 50+ more

The LLM application teams that win aren't the ones with the fanciest models. They're the ones who see what their models are doing — and iterate fastest. Langfuse is how you get that superpower. Happy building! 🪢✨

Comments