OpenTelemetry-native tracing
Record agent and LLM application activity as spans, including inputs, outputs, latency, token counts, and cost for individual steps.
TruLens is an open-source library for tracing and evaluating AI agents and other LLM applications. It helps development teams inspect agent steps, score quality with configurable metrics, and compare application versions using OpenTelemetry-based instrumentation.
TruLens is an open-source evaluation and observability library for AI agents and LLM applications. It records application activity as traces and evaluates quality at the level of an overall response or individual steps, helping teams connect an unsatisfactory answer to the retrieval, tool call, plan, or generation that caused it.
The product captures details such as inputs, outputs, latency, token usage, and cost for recorded steps. Its evaluation workflow includes metrics for agent behavior, retrieval-augmented generation, and tool use, including context relevance, groundedness, answer relevance, correctness, plan adherence, and tool selection. Evaluation results can be inspected alongside trace details rather than treated as an isolated score.
Teams can start with built-in LLM-judge metrics and adapt them with rubrics, examples, score ranges, and additional instructions. TruLens also supports custom metrics and comparison of application versions through leaderboard-style results that place quality measures alongside latency and cost, making it possible to assess changes such as revised tool arguments, reranking, policy citation, or model selection.
The documentation covers runtime and batch evaluation, evaluation of existing data, human feedback logging, ground-truth workflows, guardrails, and viewing trends and records. Instrumentation is based on OpenTelemetry, and the documentation includes guides and provider references for working with frameworks, model providers, and storage systems. TruLens is distributed under the MIT license; the supplied product pages do not provide a detailed pricing model or plan limits.
Record agent and LLM application activity as spans, including inputs, outputs, latency, token counts, and cost for individual steps.
Score agent behavior and intermediate operations with metrics such as tool selection, plan adherence, context relevance, groundedness, answer relevance, and correctness.
View score explanations that identify issues such as missing tool calls, unsuitable arguments, or poorly selected context instead of relying only on an aggregate number.
Adapt shipped evaluations with rubrics, few-shot examples, score ranges, and additional instructions, or define custom metrics with selectors for the relevant record data.
Compare application versions using evaluation scores together with average latency and total cost to investigate quality and efficiency trade-offs.
Inspect a failed answer through its trace to determine whether planning, retrieval, tool selection, tool arguments, or generation introduced the problem.
Measure whether retrieved context is relevant and whether the final answer is grounded in that context, using RAG-focused metrics and evaluation workflows.
Run comparable records across versions and compare quality, latency, and cost when changing prompts, tools, reranking, citations, or models.
Use runtime evaluation, recorded traces, result views, and trends to follow application behavior after release rather than relying only on pre-launch tests.
Refine judge behavior with domain-specific criteria, examples, and scoring instructions when general-purpose metrics do not capture the team's quality standard.
It traces AI-agent and LLM-application activity as spans. The documented trace view includes actions such as planning, retrieval, tool calls, and generation, along with inputs, outputs, timing, token usage, and cost where available.
Yes. The documented metric API supports custom criteria, additional instructions, few-shot examples, score ranges, selectors, and custom metrics. This allows teams to align evaluations with domain-specific requirements.
The documentation includes runtime evaluation and inline evaluations, as well as evaluation on existing data and batch evaluation. The exact workflow depends on the application and instrumentation setup.
No. Its examples and metrics cover intermediate agent behavior, including plan adherence, tool selection, retrieval context, and other recorded steps, in addition to final-answer quality.
No pricing tiers, limits, or commercial terms are provided in the supplied pricing-page content. The available product information identifies TruLens as open source and lists an MIT license.
www.lyzr.ai
OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.
www.langchain.com
LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.
arize.com
Phoenix is an open-source, local-first platform for tracing, evaluating, experimenting with, and improving AI applications and agents. It helps AI engineers inspect agent behavior, assess output quality, and test changes before deployment.
harproject.dev
HAR HQ is a team governance and observability layer for organizations running coding agents with HAR. It standardizes verification, tracks AI-related work and spend, and helps engineering teams review and improve agent workflows.
context.ai
Context は、顧客インフラ上で AI エージェントを構築・デプロイ・改善できる企業向けプラットフォームです。ワークスペース、ランタイム、コンテキスト、評価ツール、コネクタ、IdP ベースのアクセス制御を備えています。
www.galileo.ai
Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.