Agent tracing
View agent execution step by step, including nested runs, model calls, tool interactions, and multi-turn message threads. Traces help teams locate failures affecting latency, cost, or response quality.
LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.
LangSmith is an AI agent and LLM observability and evaluation platform. Its core purpose is to make complex agent applications inspectable: teams can trace model calls, tool use, nested runs, and multi-turn conversations, then use that information to debug failures and improve application quality.
The platform combines tracing, real-time monitoring, online and offline evaluations, dataset collection, human feedback, prompt development, and automated trace analysis. It supports applications built with LangChain as well as custom implementations and other LLM frameworks, with SDKs for Python, TypeScript, Go, and Java and support for OpenTelemetry.
LangSmith is available as a managed cloud service, with bring-your-own-cloud and self-hosted options for teams that need greater control over where trace data is stored. The product also includes SmithDB, a trace database designed for agent-oriented queries such as full-text search, JSON key-path filtering, and trajectory analysis.
View agent execution step by step, including nested runs, model calls, tool interactions, and multi-turn message threads. Traces help teams locate failures affecting latency, cost, or response quality.
Monitor token usage, P50 and P99 latency, error rates, cost breakdowns, and feedback scores through dashboards. Webhook and PagerDuty alerts can notify teams when monitored conditions require attention.
Score application behavior with online LLM-as-judge and code evaluations, and run offline evaluations against datasets. LangSmith also supports dataset collection and human annotation workflows.
Automatically analyze and cluster traces to identify usage patterns, common agent behaviors, and failure modes. Error-analysis templates and executive summaries help organize findings.
Search and debug large volumes of agent traces using random access to runs, full-text search, JSON key-path filtering, and trajectory queries. SmithDB can be self-hosted inside a VPC.
Trace applications built with popular agent and LLM frameworks, custom implementations, or existing OpenTelemetry pipelines. SDKs are available for Python, TypeScript, Go, and Java.
An engineer can inspect the complete execution path of a problematic request, including model and tool calls, to identify the step responsible for an incorrect result, excessive latency, or unexpected cost.
An operations or AI platform team can track quality and runtime signals across deployed applications, apply online evaluations, and route threshold-based alerts through webhooks or PagerDuty.
A team can collect datasets from application activity, add human feedback, and run offline evaluations to compare behavior or check whether previously observed failures recur after a prompt or implementation change.
A team investigating a large trace corpus can use Insights to cluster topics and behaviors, then use SmithDB search and trajectory queries to examine representative failures and organize error analysis.
Organizations that cannot send sensitive traces to a shared environment can consider BYOC or self-hosted deployment options, including self-hosting SmithDB inside a VPC.
It provides visibility into AI agent and LLM application behavior, including execution steps, model and tool calls, multi-turn interactions, latency, token usage, errors, costs, and feedback or evaluation scores.
Yes. The source describes support for OpenAI SDK, Anthropic SDK, Vercel AI SDK, LlamaIndex, and custom implementations, along with SDKs for Python, TypeScript, Go, and Java. OpenTelemetry can connect LangSmith with existing telemetry pipelines.
Yes. LangSmith states that Observability and Evaluation can be used independently. Teams can start with tracing and monitoring and add evaluations later.
LangSmith is offered as a managed cloud service, with bring-your-own-cloud and self-hosted options. Enterprise hosting can run on a team’s Kubernetes cluster in AWS, GCP, or Azure, while SmithDB can be self-hosted inside a VPC.
The site states that the SDK sends traces asynchronously through a distributed collector, so an incident in LangSmith should not stop the agent from running. LangSmith also states that it does not train on submitted data and that customers retain rights to their data.
www.lyzr.ai
OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.
arize.com
Phoenix is an open-source, local-first platform for tracing, evaluating, experimenting with, and improving AI applications and agents. It helps AI engineers inspect agent behavior, assess output quality, and test changes before deployment.
www.trulens.org
TruLens is an open-source library for tracing and evaluating AI agents and other LLM applications. It helps development teams inspect agent steps, score quality with configurable metrics, and compare application versions using OpenTelemetry-based instrumentation.
harproject.dev
HAR HQ is a team governance and observability layer for organizations running coding agents with HAR. It standardizes verification, tracks AI-related work and spend, and helps engineering teams review and improve agent workflows.
context.ai
Context is an enterprise AI agents platform for building, deploying, and improving agents on customer infrastructure, with workspace, runtime, context, evaluation, connectors, and IdP access controls.
www.galileo.ai
Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.