Recommendation

Choose the quality workflow, not just a trace viewer.

Define which model and agent events must be traced, which outcomes require evaluations, how datasets and feedback are governed, and whether cloud or self-hosted operation is acceptable.12

For: AI application and platform teams that need model-aware traces, evaluations, datasets, feedback, token and cost semantics, and behavior analysis

Main trade-off

AI observability improves diagnosis and evaluation while adding sensitive prompt and output capture, instrumentation, scoring cost, retention, access control, sampling, platform, and self-hosting obligations.1234

Choose the product and operating model first

Start from workload, integration, quality, policy, data, lifecycle, and ownership requirements before comparing product breadth.

  1. Telemetry model

    List required generations, prompts, retrieval, tools, handoffs, agent runs, tokens, costs, feedback, model metadata, and trace hierarchy.1

  2. Evaluation workflow

    Define datasets, expected outputs, scorers, human review, experiments, online evaluation, regression gates, and reproducibility.3

  3. Privacy and access

    Decide whether inputs, outputs, tool arguments, retrieved content, identities, and evaluator reasoning may be captured; define redaction, tenancy, retention, export, and deletion.12

  4. Operating model

    Compare cloud, self-hosted, hybrid, instrumentation, ingestion, storage, sampling, pricing units, alerting, framework coupling, and exit.12

Decision routes

Each route addresses a distinct workload or ownership boundary and retains an explicit verification condition.

Open-source LLM observability

Choose Open-source LLM observability

Choose Langfuse when open-source LLM observability, cloud or self-hosting, tracing, prompt, dataset, and evaluation workflows fit the team.

Verify: Verify current license and edition boundaries, self-host architecture, upgrades, storage, retention, auth, ingestion, pricing, export, and operational capacity.1

LangChain-aligned evaluation

Choose LangChain-aligned evaluation

Choose LangSmith when LangChain or LangGraph alignment and its managed tracing, datasets, experiments, evaluation, and deployment workflow are deliberate.

Verify: Keep LangSmith distinct from LangGraph; verify framework independence, capture controls, retention, regions, seats, traces, eval pricing, deployment, export, and policy.2

Evaluation-led development

Choose Evaluation-led development

Choose Braintrust when datasets, immutable experiments, scorers, human review, production traces, and evaluation-led iteration are the primary workflow.

Verify: Verify trace and evaluation ingestion, data capture, retention, regions, seats, scoring model cost, online evaluation, export, pricing, and hosted operating boundary.3

Open-source tracing and evaluation

Choose Open-source tracing and evaluation

Choose Phoenix when OpenTelemetry and OpenInference tracing, open-source evaluation, experiments, and self-hosting are central.

Verify: Verify Phoenix versus Arize cloud boundaries, deployment, storage, auth, retention, scaling, integrations, evaluator cost, export, and maintenance ownership.4

Boundary: AI Observability owns model, generation, retrieval, tool, agent, token, cost, prompt-version, dataset, feedback, and evaluation semantics. Logging and APM retain generic events and service telemetry; Gateway analytics cover only traffic passing through the gateway.

Differences that change the choice

Compare the workload, product boundary, policy, lifecycle, data path, cost shape, and operating ownership rather than feature volume.

Quality workflow
Tracing explains execution; evaluations measure outputs against explicit datasets and criteria. Neither substitutes for the other.2
Deployment
Cloud, open-source, self-hosted, and enterprise editions transfer storage, upgrades, identity, scale, and support differently.1
Instrumentation
Provider wrappers, framework callbacks, OpenTelemetry, and OpenInference offer different coverage and portability.4
Data boundary
Prompt, output, tool, retrieval, identity, and evaluator data require explicit capture, redaction, access, retention, and export policy.12

Official resources

Verify current model, API, SDK, product, pricing, policy, data, region, lifecycle, and operating boundaries in first-party material.

Sources

Official documentation supports current product boundaries and verification points; route selection remains a bounded editorial judgment.

  1. 1
    Langfuse official documentation

    Langfuse · Accessed Official

  2. 2
    LangSmith official documentation

    LangSmith · Accessed Official

  3. 3
    Braintrust official documentation

    Braintrust · Accessed Official

  4. 4
    Arize Phoenix official documentation

    Arize Phoenix · Accessed Official

  5. 5
    OpenTelemetry GenAI semantic conventions

    OpenTelemetry · Accessed Official

  6. 6
    Phoenix self-hosting

    Arize Phoenix · Accessed Official