9 Best LLM Observability Tools in 2026
Our Top Picks
Teams that want production-grade tracing without vendor lock-in
Teams already building on LangChain or LangGraph
Eval-driven teams that want to gate deploys on quality scores
Comparison Table
| Tool | Rating | Price | Best For | Action |
|---|---|---|---|---|
L Langfuse | 4.8 | Free self-hosted / Free cloud / $29/mo Core / $199/mo Pro | Teams that want production-grade tracing without vendor lock-in | Try Langfuse Free |
L LangSmith | 4.6 | Free Developer / $39/seat/mo Plus / Custom Enterprise | Teams already building on LangChain or LangGraph | Try LangSmith Free |
B Braintrust | 4.6 | Free Starter / $249/mo Pro / Custom Enterprise | Eval-driven teams that want to gate deploys on quality scores | Try Braintrust Free |
AP Arize Phoenix | 4.5 | Free self-hosted / AX Free / $50/mo AX Pro | RAG and retrieval-heavy apps that need deep eval primitives | Try Arize Phoenix Free |
M MLflow | 4.5 | Free (Apache 2.0, self-hosted) | Teams that want full data ownership at zero licence cost | Try MLflow Free |
H Helicone | 4.3 | Free Hobby / $79/mo Pro / $799/mo Team | Fastest possible cost and latency visibility across providers | Try Helicone Free |
WW W&B Weave | 4.2 | Free (1 GB/mo) / from $60/mo Pro | ML teams already standardised on Weights & Biases | Try W&B Weave Free |
DA Datadog Agent Observability | 4.4 | Free (40k spans/mo) / from $160/mo Pro | Enterprises correlating LLM traces with existing APM and incidents | Try Datadog Agent Observability Free |
CA Confident AI | 4.2 | Free / $200/mo Starter / $2,000/mo Team | Teams that want a managed UI on top of the DeepEval framework | Try Confident AI Free |
9 Best LLM Observability Tools in 2026
The best LLM observability tools solve a problem traditional monitoring never had to: your application can return a perfectly valid HTTP 200 and still be completely wrong. There is no stack trace for a hallucination, no exit code for a retrieval step that pulled the wrong document, and no CPU graph that explains why last Tuesday's prompt tweak quietly degraded answer quality by 15%.
That gap is why this category exploded. The LLM observability market is valued at roughly $2.69 billion in 2026 and projected to reach $9.26 billion by 2030 — a 36.2% CAGR — according to figures compiled in MarkTechPost's 2026 platform roundup. The same roundup notes that while 89% of organisations now run some form of agent observability, only 52.4% run offline evaluations — meaning most teams can see what their agents did, but cannot prove whether it was any good.
We compared the nine platforms that actually matter in 2026, with pricing and licence terms pulled from each vendor's own pages as of October 2026.
Our Top 3 Picks
- Langfuse — the best overall LLM observability tool. MIT-licensed core, free self-hosting, and a cloud tier that does not charge per seat.
- Braintrust — the best choice if you want evaluation to gate your deploys rather than just decorate a dashboard.
- Arize Phoenix — the best option for RAG-heavy applications, and the most OpenTelemetry-native of the group.
Quick Comparison
| Tool | Licence | Free tier | Paid entry | Self-host |
|---|---|---|---|---|
| Langfuse | MIT (core) | 50k units/mo, 2 users, 30-day | $29/mo Core | Yes, free |
| LangSmith | Proprietary | 5k traces/mo, 1 seat | $39/seat/mo | Enterprise only |
| Braintrust | Proprietary | 1 GB/mo, 10k scores, 14-day | $249/mo Pro | Enterprise only |
| Arize Phoenix / AX | ELv2 (source-available) | 25k spans/mo, 1 GB, 15-day | $50/mo AX Pro | Yes, free |
| MLflow | Apache 2.0 | n/a (self-host) | — | Yes, free |
| Helicone | Open source | 10k requests, 1 GB, 7-day | $79/mo Pro | Yes |
| W&B Weave | Apache 2.0 SDK | 1 GB/mo ingestion | from $60/mo Pro | Enterprise |
| Datadog | Proprietary | 40k LLM spans/mo, 15-day | from $160/mo Pro | No |
| Confident AI | Proprietary (DeepEval is OSS) | 1 project, 5 runs/week, 1 GB-mo | $200/mo Starter | Enterprise |
1. Langfuse — Best Overall
Langfuse has become the default answer for teams who want serious tracing without betting their roadmap on a vendor. The core is MIT licensed (the ee folders are excluded), the repository sits at 35.2k GitHub stars, and self-hosting via Docker Compose or Kubernetes costs nothing.
In January 2026 Langfuse was acquired by ClickHouse alongside ClickHouse's $400M Series D. The team confirmed the project stays open source and self-hostable and that Langfuse Cloud continues to run unchanged — a reassurance worth noting, because acquisition is exactly the moment open-core projects usually tighten their licence.
Feature-wise you get nested trace visualisation, LLM-as-judge evaluators, human annotation queues, prompt management, and dataset regression testing you can wire into GitHub Actions. Cloud pricing runs Hobby (free, 50k units/month, 2 users, 30-day access), Core at $29/mo, Pro at $199/mo with three-year data access, and Enterprise at $2,499/mo. Overage beyond the included 100k units starts at $8.00 per 100k and tapers to $6.00 at high volume.
Verdict: pick Langfuse unless you have a specific reason not to. The one wrinkle is that "units" are not the same as traces, so model your real volume before committing to a tier.
2. Braintrust — Best for Eval-Driven Development
Braintrust inverts the usual priority. Tracing exists, but the centre of gravity is the experiment framework: define a dataset, run prompt or model variations against it, and compare results side by side. Combined with CI regression testing, this is the cleanest path to gating a deployment on eval scores rather than merging on vibes.
The free Starter tier is unusually usable — unlimited users, projects, datasets and experiments, plus $10/month of model credits, 1 GB of processed data, and 10k scores at 14-day retention. Pro at $249/mo lifts that to $100 in credits, 5 GB, 50k scores and 30-day retention, and adds RBAC, custom charts and environments. Qualifying startups can get 6–12 months free.
Verdict: the best tool here for teams who treat prompts as code. The $0 → $249/mo cliff is the main objection; there is nothing in between.
3. Arize Phoenix — Best for RAG and OpenTelemetry
Phoenix is built on top of OpenTelemetry, not merely compatible with it. The arize-phoenix-otel packages give you OTel primitives with sane defaults, which means traces can land in Phoenix and your existing OTLP backend simultaneously — the strongest anti-lock-in story in this list.
For retrieval-augmented applications it is the clear leader: retrieval relevance scoring, embedding drift detection, document-level attribution, and RAG-specific quality plots. It also ships a remote MCP server, so Claude Code or Cursor can query your traces directly. pip install arize-phoenix then phoenix serve is the whole setup.
The licence deserves care. Phoenix uses Elastic License 2.0 — source-available, 11.7k stars, free to read, modify and self-host, but you may not offer it as a hosted service to third parties. The managed AX Free tier allows 25k spans/month, 1 GB and 15-day retention; AX Pro at $50/mo raises that to 50k spans, 10 GB and 30 days. Evals and users are unlimited on every tier.
Verdict: best technical foundation of the group. Just do not call it open source in a compliance review.
4. LangSmith — Best for LangChain and LangGraph Stacks
If you are already on LangChain 1.0 or LangGraph 1.0, LangSmith is the default backend and the integration cost is close to zero. You get full conversation traces, the Polly assistant for trace summarisation, issue clustering via the LangSmith Engine, calibrated judges, and unified cost tracking across agent workflows.
Pricing is Developer (free, 5k base traces/month, 1 seat), Plus at $39/seat/month with 10k base traces, unlimited seats and one free small serverless deployment, and Enterprise with custom pricing plus cloud, hybrid or self-hosted deployment, SSO and RBAC. Usage beyond the base is billed in LangChain Compute Units at $1.50/LCU and Storage Units at $1.00/LSU. Base traces retain for 14 days; extended retention to 180 days costs extra.
Verdict: excellent inside the LangChain ecosystem, hard to justify outside it. Stacking per-seat fees on top of usage charges is the least predictable pricing model here.
5. MLflow — Best Free and Fully Owned
MLflow 3 added genuine GenAI tracing to a tool most ML teams already run. It is Apache 2.0 under the Linux Foundation, 100% free, and there is no commercial tier you eventually get pushed toward — trace data stays entirely yours.
You get native agent tracing, OTel GenAI semantic conventions for both export and ingestion, LLM judges, and prompt optimisation using GEPA and MIPRO. Auto-instrumentation covers OpenAI, LangChain, LlamaIndex, DSPy and Pydantic AI, and it integrates with RAGAS, DeepEval and TruLens for evaluation.
Verdict: the right answer when procurement or data residency rules out SaaS entirely. The trade-off is that you own the uptime, scaling and upgrades.
6. Helicone — Best for Fast Cost Visibility
Helicone is the lowest-effort option in this list. It is a proxy: change your base URL and you immediately have request logs, token counts and multi-provider cost breakdowns, with response caching and prompt experimentation on top. No SDK instrumentation, no decorators.
Hobby is free with 10k requests, 1 GB storage, 1 seat and 7-day retention. Pro at $79/mo adds unlimited seats, alerts, reports, the HQL query language, one-month retention and 1,000 logs/minute ingestion. Team at $799/mo brings SOC-2 and HIPAA compliance, three-month retention and 15,000 logs/minute. Enterprise adds SAML SSO, on-premises deployment and 30,000 logs/minute.
Verdict: best first instrument when you need a cost answer this afternoon. The request-centric design gives shallower insight into multi-step agent reasoning, and a proxy in your request path is a real architectural decision.
7. Datadog Agent Observability — Best for Enterprise Correlation
Datadog's LLM offering is now folded into Agent Observability, and its advantage is entirely about context. Nowhere else can you pivot from a degraded LLM span to the database latency, deploy event and on-call incident that explain it. Prompt-injection and other security signals are built in, and it inherits Datadog's 1,000+ integrations.
The free tier is genuinely generous: up to 40k LLM spans/month with 15-day retention, unlimited context and unlimited evaluations. Pro from $160/mo (annual) includes 100k LLM spans, with overage at $3.50 per 10k spans annually, $4.20 month-to-month, or $5.00 on demand. Extended retention is an add-on at $1.50–$4.00 per 10k spans depending on window.
Verdict: the obvious pick if Datadog is already your observability backbone. If it is not, the per-span economics and SaaS-only deployment make it an expensive entry point, and its offline eval tooling is the weakest here.
8. W&B Weave — Best for Existing Weights & Biases Teams
Weave extends W&B into LLM territory with structured execution traces for multi-agent systems and lineage that connects back to your existing experiment tracking. The SDK is Apache 2.0; the cloud is commercial.
The Free plan includes 1 GB/month of Weave ingestion. Pro starts at $60/month with 1.5 GB included and additional ingestion at $0.10/MB. There is also a free-forever academic licence covering 25 GB/month and up to 100 seats, which is the most generous research offer in this comparison.
Verdict: a sensible extension if you are already a W&B shop. Note that ingestion-based pricing is unkind to large prompts and long context windows — a single verbose agent run can consume a surprising slice of your monthly GB.
9. Confident AI — Best Managed Layer for DeepEval
Confident AI is the commercial platform built around DeepEval, its widely adopted open-source evaluation framework. If your team already writes DeepEval test cases, this adds a hosted UI, no-code eval workflows, custom metrics, online evals, annotation queues and real-time alerting.
Free covers 2 seats, 1 project, 5 test runs per week and 1 GB-month of trace spans. Starter at $200/mo unlocks the no-code workflows with 5 projects and 5 GB-months included, then $1/GB-month. Team at $2,000/mo adds metric versioning, git-based prompt workflows, custom RBAC, SOC2 and SSO with 75 GB-months. Enterprise adds on-prem, HIPAA and 24/7 support.
Verdict: worth it specifically as a DeepEval accelerator. As a general observability platform the tracing depth does not match Langfuse or Phoenix, and the free tier's 5-runs-per-week cap is restrictive.
How to Choose
Decide in this order:
- Does SaaS pass compliance? If not, your shortlist is MLflow, Langfuse self-hosted, or Phoenix self-hosted.
- Is your bottleneck debugging or quality? Debugging production incidents points to Langfuse, Phoenix or Datadog. Proving quality before release points to Braintrust or Confident AI.
- Are you locked to a framework or vendor? LangChain users should default to LangSmith; W&B users to Weave; Datadog users to Agent Observability.
- Do you just need a cost number? Helicone, this afternoon.
- Is retrieval your weak point? Phoenix, for the RAG-specific diagnostics nothing else matches.
One practical warning on budgeting: these platforms meter in four incompatible units — traces, spans, units and gigabytes of ingestion. A 10k-request/day agent that makes six LLM calls per request is 1.8M spans a month, which lands in a very different price bracket on Datadog than on Langfuse. Instrument one week of real traffic on a free tier before you sign anything annual.
Frequently Asked Questions
What is LLM observability? Capturing the inputs, outputs, intermediate steps, latency, token usage and cost of every LLM or agent call, then evaluating output quality. Conventional APM tells you a request succeeded; LLM observability tells you whether the answer was correct.
What is the difference between observability and evaluation? Observability is reactive and online — what happened in production. Evaluation is proactive and usually offline — scoring outputs against a dataset before release. Langfuse, Phoenix and Datadog lead on the first; Braintrust and Confident AI lead on the second. Most teams need both.
Which LLM observability tool is genuinely open source?
MLflow (Apache 2.0), Langfuse's core (MIT, excluding ee folders) and Helicone. Arize Phoenix is source-available under Elastic License 2.0, which permits self-hosting but forbids reselling it as a service — not OSI-approved open source.
Can I avoid vendor lock-in? Prefer OpenTelemetry-native tools. Phoenix is built on OTel and MLflow supports GenAI semantic conventions for both export and ingestion, so instrumentation written once can be pointed at a different backend later.
Do I need a paid plan to start? No. Datadog (40k spans/month), Langfuse (50k units/month), Braintrust (unlimited users, 1 GB) and Arize AX (25k spans/month) all have free tiers large enough for a real pilot, and MLflow is free at any volume if you self-host.
Pricing and licence terms verified against vendor pages in October 2026. LLM platform pricing changes frequently — confirm current figures before purchasing.
Sources
Pros
- MIT-licensed core and genuinely free self-hosting
- Framework-agnostic SDKs plus prompt management and LLM-as-judge evals
- Unlimited users from the $29/mo tier upward
Cons
- Cloud overages are metered in 'units' that are easy to misjudge
- Enterprise tier jumps to $2,499/mo
- Self-hosting the full stack means running ClickHouse yourself
Pros
- Default, zero-config backend for LangChain 1.0 and LangGraph 1.0
- Strong agent trace visualisation and issue clustering
- Hybrid and self-hosted deployment on Enterprise
Cons
- Proprietary — no open-source option
- Per-seat pricing on top of usage-based LCU/LSU charges
- Base traces retained only 14 days unless you pay to extend
Pros
- Unlimited users, projects, and experiments on every tier including free
- Best-in-class dataset versioning and CI regression testing
- Includes monthly model credits ($10 free, $100 on Pro)
Cons
- Big jump from free to $249/mo with no tier between
- 14-day retention on the free tier
- Tracing is solid but secondary to the eval workflow
Pros
- Built directly on OpenTelemetry — exports to any OTLP backend
- Strongest RAG debugging: retrieval relevance, drift, document attribution
- Runs locally with two commands and ships a remote MCP server
Cons
- Elastic License 2.0 is source-available, not OSI open source
- AX Free caps at 25k spans/month and 15-day retention
- Two overlapping products (Phoenix and AX) confuse onboarding
Pros
- Apache 2.0 under the Linux Foundation — no commercial tier required
- Native GenAI tracing plus OTel GenAI semantic conventions
- Prompt optimisation and integrations with RAGAS, DeepEval and TruLens
Cons
- You operate and scale the backend yourself
- UI is less polished than the commercial dashboards
- No managed free tier to prototype against
Pros
- One-line proxy integration — no SDK instrumentation needed
- Response caching and multi-provider cost breakdown out of the box
- Open source and self-hostable
Cons
- Proxy sits in your request path, adding a failure point
- 7-day retention and 10k requests on the free tier
- Request-centric model gives shallower multi-step agent traces
Pros
- Shares lineage and experiment tracking with existing W&B projects
- Clear structured traces for multi-agent systems
- Free-forever academic licence includes 25 GB/mo and up to 100 seats
Cons
- Ingestion-based pricing punishes long prompts and large payloads
- Overage at $0.10/MB adds up quickly
- Little value unless you are already a W&B customer
Pros
- Correlates AI spans with infrastructure, APM and incident workflows
- Prompt-injection and security signals built in
- Generous 40k-span free tier to evaluate with
Cons
- SaaS only — no self-hosting
- Per-10k-span overages and paid retention add-ons compound
- Weakest offline evaluation tooling of this group
Pros
- Built on DeepEval, a widely used open-source eval framework
- No-code eval workflows, annotation queues and real-time alerting
- Transparent $1/GB-month overage on trace spans
Cons
- Free tier limited to 1 project and 5 test runs per week
- Starter at $200/mo is steep for small teams
- Observability is thinner than dedicated tracing platforms