TruLens logo

TruLens

Freemium
Visit

TruLens is an open-source library for tracing and evaluating AI agents and other LLM applications. It helps development teams inspect agent steps, score quality with configurable metrics, and compare application versions using OpenTelemetry-based instrumentation.

What is TruLens?

TruLens is an open-source evaluation and observability library for AI agents and LLM applications. It records application activity as traces and evaluates quality at the level of an overall response or individual steps, helping teams connect an unsatisfactory answer to the retrieval, tool call, plan, or generation that caused it.

The product captures details such as inputs, outputs, latency, token usage, and cost for recorded steps. Its evaluation workflow includes metrics for agent behavior, retrieval-augmented generation, and tool use, including context relevance, groundedness, answer relevance, correctness, plan adherence, and tool selection. Evaluation results can be inspected alongside trace details rather than treated as an isolated score.

Teams can start with built-in LLM-judge metrics and adapt them with rubrics, examples, score ranges, and additional instructions. TruLens also supports custom metrics and comparison of application versions through leaderboard-style results that place quality measures alongside latency and cost, making it possible to assess changes such as revised tool arguments, reranking, policy citation, or model selection.

The documentation covers runtime and batch evaluation, evaluation of existing data, human feedback logging, ground-truth workflows, guardrails, and viewing trends and records. Instrumentation is based on OpenTelemetry, and the documentation includes guides and provider references for working with frameworks, model providers, and storage systems. TruLens is distributed under the MIT license; the supplied product pages do not provide a detailed pricing model or plan limits.

What can TruLens do?

OpenTelemetry-native tracing

Record agent and LLM application activity as spans, including inputs, outputs, latency, token counts, and cost for individual steps.

Step-level evaluation

Score agent behavior and intermediate operations with metrics such as tool selection, plan adherence, context relevance, groundedness, answer relevance, and correctness.

Explainable LLM judges

View score explanations that identify issues such as missing tool calls, unsuitable arguments, or poorly selected context instead of relying only on an aggregate number.

Configurable and custom metrics

Adapt shipped evaluations with rubrics, few-shot examples, score ranges, and additional instructions, or define custom metrics with selectors for the relevant record data.

Version and quality-cost comparison

Compare application versions using evaluation scores together with average latency and total cost to investigate quality and efficiency trade-offs.

Use Cases

“Debugging agent failures”

Inspect a failed answer through its trace to determine whether planning, retrieval, tool selection, tool arguments, or generation introduced the problem.

“Evaluating RAG pipelines”

Measure whether retrieved context is relevant and whether the final answer is grounded in that context, using RAG-focused metrics and evaluation workflows.

“Testing application changes”

Run comparable records across versions and compare quality, latency, and cost when changing prompts, tools, reranking, citations, or models.

“Monitoring deployed applications”

Use runtime evaluation, recorded traces, result views, and trends to follow application behavior after release rather than relying only on pre-launch tests.

“Aligning evaluations to a domain”

Refine judge behavior with domain-specific criteria, examples, and scoring instructions when general-purpose metrics do not capture the team's quality standard.

Frequently Asked Questions

What does TruLens trace?

It traces AI-agent and LLM-application activity as spans. The documented trace view includes actions such as planning, retrieval, tool calls, and generation, along with inputs, outputs, timing, token usage, and cost where available.

Can evaluation scores be customized?

Yes. The documented metric API supports custom criteria, additional instructions, few-shot examples, score ranges, selectors, and custom metrics. This allows teams to align evaluations with domain-specific requirements.

Can TruLens evaluate an application while it runs?

The documentation includes runtime evaluation and inline evaluations, as well as evaluation on existing data and batch evaluation. The exact workflow depends on the application and instrumentation setup.

Does TruLens only evaluate final answers?

No. Its examples and metrics cover intermediate agent behavior, including plan adherence, tool selection, retrieval context, and other recorded steps, in addition to final-answer quality.

Is detailed pricing available on the supplied product pages?

No pricing tiers, limits, or commercial terms are provided in the supplied pricing-page content. The available product information identifies TruLens as open source and lists an MIT license.

Quick Facts

Category
Developer Tool
Primary function
AI-agent tracing and evaluation
Instrumentation
OpenTelemetry-native
License
MIT
Evaluation modes
Runtime, batch, existing-data, and ground-truth workflows
Source domain
trulens.org

TruLens Alternatives

OpenController logo

OpenController

www.lyzr.ai

OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.

LangSmith logo

LangSmith

www.langchain.com

LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.

Phoenix logo

Phoenix

arize.com

Phoenix is an open-source, local-first platform for tracing, evaluating, experimenting with, and improving AI applications and agents. It helps AI engineers inspect agent behavior, assess output quality, and test changes before deployment.

HAR HQ logo

HAR HQ

harproject.dev

HAR HQ is a team governance and observability layer for organizations running coding agents with HAR. It standardizes verification, tracks AI-related work and spend, and helps engineering teams review and improve agent workflows.

Context logo

Context

context.ai

Context is an enterprise AI agents platform for building, deploying, and improving agents on customer infrastructure, with workspace, runtime, context, evaluation, connectors, and IdP access controls.

Galileo logo

Galileo

www.galileo.ai

Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.