Pytest-native evaluation
Write LLM tests as Python tests, run them locally or in CI/CD, and invoke them through the `deepeval` CLI.
DeepEval is an open-source LLM evaluation framework for testing and benchmarking AI applications. It helps developers run pytest-native evaluations, score outputs and agent traces, and iterate on systems across text, image, audio, and voice workflows.
DeepEval is an open-source evaluation framework for LLM applications. It provides pytest-native tests and Python-based evaluation flows that can run locally, in CI/CD, or through the deepeval CLI. Teams can evaluate black-box outputs, complete agent trajectories, and individual component decisions using configurable metrics and test cases.
The framework includes more than 50 ready-to-use metrics for areas such as hallucination, faithfulness, answer relevancy, summarization, toxicity, bias, agent task completion, tool correctness, retrieval quality, and multi-turn conversation behavior. Metrics generally return a score from 0 to 1 together with reasoning, and a test passes when its score meets the configured threshold. DeepEval also supports custom criteria through techniques including G-Eval, JevEval, DAG metrics, and self-coded metrics.
DeepEval covers text, images, and audio in the same evaluation model. It can trace agent execution so teams can inspect and grade ordered steps, tools, retrievers, and model calls. Integrations are available for agent frameworks, model providers, voice systems, vector databases, and common CI/CD platforms. For shared reports and broader evaluation operations, DeepEval connects with the Confident AI platform.
Write LLM tests as Python tests, run them locally or in CI/CD, and invoke them through the `deepeval` CLI.
Use ready-made metrics for agents, RAG, chatbots, safety, multimodal evaluation, hallucination, JSON correctness, summarization, and related criteria.
Capture ordered agent execution and evaluate whole trajectories or individual actions such as tool selection and argument construction.
Define natural-language criteria with G-Eval or bounded questions with JevEval, build conditional DAG metrics, or implement self-coded metrics.
Generate synthetic golden test cases from sources such as knowledge-base files and simulate conversations across user personas before deployment.
Connect with frameworks including LangChain, LlamaIndex, CrewAI, LangGraph, OpenAI Agents, Pydantic AI, and others; provider, speech, vector database, and CI/CD integrations are also documented.
Add evaluation cases to a Python test suite and rerun them in a CI/CD workflow to identify quality regressions before shipping changes.
Inspect a complete agent trace, then score planning, task completion, step efficiency, tool correctness, or tool arguments to locate a failing action.
Assess retriever context relevancy, precision, and recall separately from generator answer relevancy and faithfulness.
Evaluate multi-turn behavior such as knowledge retention, role adherence, conversation completeness, and relevancy; voice workflows can use speech and voice-agent integrations.
Generate goldens from documentation or simulate user conversations and edge cases when a manually labeled dataset is not yet available.
It is primarily a code-first evaluation framework. Tests can run in Python, through pytest, or with the `deepeval` CLI; results can also be sent to Confident AI for shared reports and broader evaluation workflows.
The documented metrics and workflows cover LLM applications, AI agents, RAG systems, chatbots, multimodal applications, and voice agents. Evaluations can target final outputs, complete traces, or individual components.
Most metrics produce a score between 0 and 1 and include reasoning for the score. A metric is successful when its score reaches the configured threshold, which defaults to 0.5 according to the metrics documentation.
Yes. DeepEval supports custom criteria through G-Eval, JevEval, DAG metrics, and fully self-coded metrics, including metrics based on measures such as BLEU or ROUGE.
The supplied pricing URL currently returns a 404 page, so pricing, plan limits, and commercial terms cannot be confirmed from the available source material. The framework itself is described as open source.
www.galileo.ai
Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.
jev-state.vercel.app
Jev State is a workspace for defining, testing, and regression-checking conversational decisions powered by Jev. It helps teams inspect why an agent takes a step and export runnable TypeScript and tests for an application.
www.giskard.ai
Giskard is an AI security and evaluation platform for testing conversational LLM agents before and after deployment. It combines automated red teaming, quality evaluation, runtime guardrails, and remediation workflows for teams responsible for reliable AI systems.
www.promptfoo.dev
Promptfoo is an AI security and testing platform for evaluating LLM applications, agents, models, and workflows. It helps developers and security teams find vulnerabilities, validate guardrails, map findings to security frameworks, and track remediation through development and deployment.
www.mcpjam.com
MCPJam is a testing and evaluation platform for MCP servers. It helps developers inspect servers locally, run user and model-based tests, and add behavior checks to CI/CD workflows.
qagent.in
QAgent is an AI agent testing and quality assurance platform for developers and agile teams. It connects to an agent through a webhook or endpoint, runs automated test cases, and evaluates responses for groundedness, policy adherence, prompt compliance, and related quality dimensions.