Oqoqo logo

Oqoqo

Freemium
訪問

Oqoqo is an evaluation platform for running repeatable agent experiments in realistic, isolated environments. Teams use it to build private benchmarks, compare agents and models, and investigate why agent workflows succeed or fail.

Oqoqoとは?

Oqoqo is a platform for designing, running, and analyzing evaluations of AI agents on real-world tasks. An experiment combines task definitions and rubrics with agents, models, effort settings, product treatments, and repeated trials, allowing teams to compare behavior under controlled conditions.

Each trial runs in a fresh sandbox containing the project files, credentials, dependencies, and tools required by the task. Oqoqo records the full trajectory and reports both outcome measures, such as pass rate and lift, and operational measures, including steps, tool calls, tokens, and duration.

The platform is intended for teams building private benchmarks, evaluating agent-facing products, or investigating workflow failures. Findings include evidence-backed frictions and can inform a fix-and-rerun cycle, while results remain limited to the tasks, agents, models, and environments selected for the experiment.

Oqoqoでできること

Cross-product experiment configuration

Define tasks, rubric requirements, agents, treatments, and repeated runs in one experiment. Treatments can represent packages of skills, MCP servers, CLIs, or SDKs, with one setup used as the baseline for comparison.

Fresh sandboxed trial environments

Run each trial in an isolated, non-reused environment with project files, credentials, dependencies, and installed tools. Separate setup, agent work, and teardown timings help distinguish infrastructure time from agent time.

Full trajectory capture

Review the steps, tool calls, commands, outputs, and files changed during a run. The recorded trajectory shows where an agent stopped or took an unproductive path, rather than reducing the result to a single score.

Requirement-level evaluation

Grade every rubric requirement with a pass or fail result, a reason, and a citation into the run. Recorded work can be re-evaluated after a rubric change without running the agent again.

Comparative metrics and friction analysis

Compare pass rates and lift for treatments and agents alongside steps, tokens, tool calls, and duration. Grouped friction findings describe the obstacle, severity, recovery, likely owner, and supporting trajectory evidence.

Web, CLI, and MCP access

People can launch experiments from the web app, while agents can launch them through CLI or MCP access. Tasks and experiment configurations are versioned, and launching snapshots the treatments, skills, and environment used.

利用シーン

“Build a private benchmark for a product workflow”

Turn real-world workflows into versioned task sets with explicit rubrics. Repeated trials provide a team-owned benchmark for the workflows represented by those tasks.

“Compare agent and model combinations”

Run the same task, treatment, and trial grid across different agents, models, and effort settings. Use pass rates, lift, tokens, and duration to compare outcome and operating cost together.

“Evaluate agent-facing interfaces and tools”

Test whether agents can use a product setup that includes skills, MCP servers, CLIs, or SDKs. Compare a baseline treatment with an improved setup to measure the effect of the change.

“Diagnose and fix failed runs”

Inspect the trajectory and requirement citations to locate the failure, then use friction evidence to distinguish product, documentation, environment, harness, agent, and task issues before rerunning.

“Trigger evaluations from development workflows”

Use the CLI or MCP access to launch experiments when a change may affect agent workflows. Versioned tasks and snapshotted treatments help keep the measured configuration identifiable.

よくある質問

What is a run in Oqoqo?

A run is one agent attempt on one task under one treatment. Errored or cancelled runs do not count against the organization’s run pool. Model inference is separate from Oqoqo’s platform and infrastructure fees.

How does Oqoqo evaluate a task?

An evaluator reviews the recorded trajectory and changed files requirement by requirement. Each requirement receives a pass or fail result with a reason and a citation to the run; the task passes only when all requirements pass.

Can a rubric be changed without rerunning the agent?

Yes. Recorded work remains fixed, so a revised rubric can be applied through a re-evaluation of the same run. If the task itself changes, Oqoqo’s methodology says to launch a new experiment.

What does Oqoqo measure besides pass rate?

Reports include lift against a baseline, steps, tool calls, tokens, and duration. Friction analysis identifies specific obstacles and records their family, severity, recovery, likely owner, and supporting trajectory evidence.

What are the main limitations of the results?

Results describe only the selected tasks and do not establish performance across every workflow. They also depend on the agents, models, and effort settings used, which can change over time. Evaluator judgments may be incorrect, so verdict reasons and citations should be checked.

クイック情報

Category
AI evaluation and developer tool
Primary users
Teams building or testing agent-facing products and workflows
Access
Web app, CLI, and MCP
Execution model
Repeated agent trials in fresh, isolated cloud sandboxes
Free plan
100 plan runs per month; unused plan runs do not roll over
Billing model
Oqoqo runs cover platform and cloud infrastructure fees; model inference is separate

Oqoqoの代替品

OpenTrain AI logo

OpenTrain AI

opentrain.ai

OpenTrain AIは、事前審査済みのAIトレーナー、データラベラー、ドメイン専門家を採用できるマーケットプレイス兼マネージドサービスです。RLHF、評価、レッドチーム、アノテーション、エージェントワークフローに対応します。

OpenController logo

OpenController

www.lyzr.ai

OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.

ComputeArena logo

ComputeArena

computearena.ai

ComputeArena is a community benchmark directory for measuring local AI model throughput across chips, quantisations, and runtimes. It helps developers run offline benchmarks, inspect signed reports, and compare decode and prefill performance on compatible workloads.

Context logo

Context

context.ai

Context は、顧客インフラ上で AI エージェントを構築・デプロイ・改善できる企業向けプラットフォームです。ワークスペース、ランタイム、コンテキスト、評価ツール、コネクタ、IdP ベースのアクセス制御を備えています。

LangSmith logo

LangSmith

www.langchain.com

LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.

DeepEval logo

DeepEval

deepeval.com

DeepEval is an open-source LLM evaluation framework for testing and benchmarking AI applications. It helps developers run pytest-native evaluations, score outputs and agent traces, and iterate on systems across text, image, audio, and voice workflows.