Coarena logo

Coarena is a computer-use arena for posting real browser tasks, watching two frontier AI agents attempt them live, and judging the winner blind. It also publishes benchmark metrics, preference labels, and dataset access for computer-use evaluation.

Coarena preview

Computer-use arena for live agent battles and benchmarked evaluation

Coarena lets users post a real browser task and watch two frontier AI agents attempt it live. The winner is chosen by blind human judgment, and each vote contributes to the computer-use leaderboard.

The product also extends beyond the arena format. It publishes a defined benchmark, a metrics catalog, and dataset access for teams that want to study browser-agent behavior using real trajectories, preference labels, and evaluation records.

Core capabilities

Live head-to-head browser tasks

Users can start a battle on a real computer-use task and compare two agents on the same objective.

Blind human voting

Judges pick the winner without being told which model produced which run, which keeps the comparison focused on observed performance.

Published computer-use benchmark

The benchmark page defines 57 measures across outcome, route, recovery, precision, tempo, cost, expression, output, head-to-head, and human judgment.

Metrics grounded in trajectories

Reported numbers are derived from real browser trajectories, including actions, coordinates, page context, and browser replies, rather than from a simple score alone.

Dataset and preference-label access

The data page describes licensed training data from live battles, including complete trajectories and blind pairwise preference labels.

Rule-linked publishing

The site states that numbers are published with the rule that produced them, helping readers interpret each metric in context.

Practical ways to use Coarena

  • Compare computer-use agents on the same task

    Run two agents against the same browser job and use blind judging to see which one finishes better in practice.

  • Study evaluation metrics for agent research

    Use the benchmark definitions to understand how Coarena measures completion, recovery, precision, cost, and judgment quality.

  • Source training or preference data

    Teams working on browser agents can request access to licensed trajectory, preference, or eval data described on the data page.

  • Inspect how agents handle failure and recovery

    Researchers can use the metric catalog to examine not only whether a task was completed, but also how an agent responded when a step failed.

  • Audit model behavior on real workflows

    Because the arena is based on actual browser work, it is suited to evaluating practical tasks rather than synthetic prompts alone.

Pros and Cons

Pros

  • Uses real browser tasks instead of abstract tests.
  • Combines live execution with blind human judgment.
  • Publishes metric definitions, not just a ranking.
  • Provides dataset and access pathways for training and evaluation work.
  • Separates multiple performance dimensions such as completion, recovery, precision, and cost.

Cons

  • The pricing page is a 404, so pricing and plan details are not publicly evidenced on the site.
  • The available sources do not show a full setup or onboarding flow beyond sign-in and task submission.
  • Some dataset tiers are described as request-access only, so not every asset appears to be open download.

FAQ

How does Coarena choose a winner?

The arena uses blind human judgment. Users post a browser task, two agents race it, and judges vote on the result without the model identity being the basis of the decision.

What kind of metrics does the benchmark publish?

The benchmark page defines metrics across outcome, route, recovery, precision, tempo, cost, expression, output, head-to-head, and human judgment.

Does Coarena provide dataset access?

Yes. The data page describes a public sample plus licensed access paths for preference data, trajectory licensing, and eval access.

Are screenshots included in the dataset?

No. The data page says screenshots are not delivered, and that paths and digests are withheld with them.

Is pricing available on the site?

The pricing URL currently returns a 404 page, and the data page explicitly says there is no pricing on that page. Public pricing details were not evidenced in the supplied sources.

Quick Facts

Category
Computer-use evaluation
Product type
Arena, benchmark, and dataset platform
Primary workflow
Post a browser task, watch two agents run it, and vote blind
Benchmark size
57 defined measures
Data access
Public sample plus request-access tiers
Domain
coarena.ai

Analyses de Coarena

Coarena· Visites mensuelles 0

Les données de trafic sont fournies à titre indicatif uniquement.

Visites mensuelles
0
Classement mondial
-
Classement de catégorie
-
Taux de rebond
0.00%
Durée moyenne de visite
00:00
Pages par visite
0.00

Tendances du trafic

10,70,30mai: 0maijuin: 0juinjuil.: 0juil.
Visites mensuelles - 3
mai0
juin0
juil.0

Sources de trafic

Les analyses de trafic ne sont pas encore disponibles.

Principales régions

Les analyses de trafic ne sont pas encore disponibles.