LLM Evaluation

发现大语言模型评测工具、基准测试与平台,衡量模型质量、对比输出、识别风险并持续改进 AI 应用。

AI 基础设施

探索此分类集合

产品

Artificial Analysis preview
Artificial Analysis logo

Artificial Analysis

LLM Evaluation

用于比较模型、提供商和硬件的 AI 基准测试与洞察平台,提供 Pro、Enterprise 计划及免费 API。

Weights & Biases preview
Weights & Biases logo

Weights & Biases

LLM Evaluation

用于实验跟踪、制品管理、GenAI 追踪与评估的机器学习和 AI 平台。

Arena preview
Arena logo

Arena

AI聊天机器人

Arena 是一个公开的 AI 模型对比与排名平台,支持与模型聊天、为输出结果投票,并浏览按任务分类的排行榜,覆盖文本、视觉、代码、搜索、视频和智能体任务。

Pydantic preview
Pydantic logo

Pydantic

AI Observability

用于构建类型安全应用、观测 AI 系统、评估输出和路由模型访问的 AI 工程技术栈

Scale AI preview
Scale AI logo

Scale AI

LLM Evaluation

为构建可靠生产级 AI 系统的团队提供数据、评测和企业级 AI 工具

Langfuse preview
Langfuse logo

Langfuse

AI Observability

开源 AI 工程平台,用于追踪、评估和改进 LLM 应用与智能体。

SuperAnnotate preview
SuperAnnotate logo

SuperAnnotate

LLM Evaluation

面向图像、视频、文本和音频工作流的 AI 数据标注与评估管线平台。

Horizon preview
Horizon logo

Horizon

LLM Evaluation

Labelbox 的 RL 环境产品,用于后训练和评估长时程 AI 任务。

Outlier AI preview
Outlier AI logo

Outlier AI

LLM Evaluation

远程 AI 训练平台,连接专家参与灵活项目并按周获得报酬

Context logo
Context logo

Context

AI Agent Builder

Context 是企业级 AI 智能体平台,可在客户基础设施上构建、部署和改进智能体,并提供工作区、运行时、上下文、评估工具、连接器及基于 IdP 的访问控制。

Klavis AI logo
Klavis AI logo

Klavis AI

AI Agent Infrastructure

Klavis AI 提供用于训练和评估 AI 智能体的实时环境,支持编程数据和工具使用数据。

Agnost AI preview
Agnost AI logo

Agnost AI

AI Observability

Agnost AI monitors AI agents in production by analyzing conversations, traces, user corrections, and frustration. It helps teams identify recurring failures, review suggested prompt fixes, and apply approved changes with replay-based regression coverage.

OpenTrain AI preview
OpenTrain AI logo

OpenTrain AI

LLM Evaluation

OpenTrain AI 是一个市场与托管服务平台,可招聘经过预审的 AI 训练师、数据标注员和领域专家,用于 RLHF、评估、红队测试、标注及智能体工作流。

QAgent preview
QAgent logo

QAgent

AI Testing Assistant

QAgent is an AI agent testing and quality assurance platform for developers and agile teams. It connects to an agent through a webhook or endpoint, runs automated test cases, and evaluates responses for groundedness, policy adherence, prompt compliance, and related quality dimensions.

inferock-bench preview
inferock-bench logo

inferock-bench

AI Gateway And Routing

inferock-bench is a local diagnostic proxy for tracking LLM API usage, provider-reported costs, failures, and billing-integrity signals. It helps developers inspect calls to OpenAI, Anthropic, Gemini Developer API, and pinned OpenRouter endpoints using locally stored receipts.

NovaSynth preview
NovaSynth logo

NovaSynth

AI Testing Assistant

NovaSynth is a pre-production testing platform for AI agents that uses synthetic users across real SIP phone calls, LiveKit audio, and chat. It helps teams rehearse multi-turn interactions, surface failures, and validate fixes before releasing agents to customers.

MCPJam preview
MCPJam logo

MCPJam

AI Testing Assistant

MCPJam is a testing and evaluation platform for MCP servers. It helps developers inspect servers locally, run user and model-based tests, and add behavior checks to CI/CD workflows.

Oqoqo preview
Oqoqo logo

Oqoqo

LLM Evaluation

Oqoqo is an evaluation platform for running repeatable agent experiments in realistic, isolated environments. Teams use it to build private benchmarks, compare agents and models, and investigate why agent workflows succeed or fail.

ComputeArena preview
ComputeArena logo

ComputeArena

LLM Evaluation

ComputeArena is a community benchmark directory for measuring local AI model throughput across chips, quantisations, and runtimes. It helps developers run offline benchmarks, inspect signed reports, and compare decode and prefill performance on compatible workloads.

Jev State preview
Jev State logo

Jev State

AI Agent Builder

Jev State is a workspace for defining, testing, and regression-checking conversational decisions powered by Jev. It helps teams inspect why an agent takes a step and export runnable TypeScript and tests for an application.