Arena is a public AI model leaderboard and comparison platform for chatting, comparing responses, and voting across text, vision, image, video, search, and agent tasks.

Arena preview

What Arena is

Arena is a public AI model leaderboard and comparison platform built around real-user evaluation. It lets people chat with multiple models, compare their answers, and vote on results that help shape rankings for large language models, image systems, code models, and agentic tools.

The site organizes model performance into dedicated arenas, including text, vision, web development, search, video, image editing, and agent tasks. That makes it useful both for casual comparison and for readers who want a quick view of how frontier models are performing across different kinds of work.

Core features

Chat and compare models

Run side-by-side chats and compare model responses in one interface, with voting that feeds the community ranking process.

Battle-style evaluation

Use Battle Mode to evaluate frontier models through direct head-to-head interactions instead of relying only on static benchmark tables.

Multi-arena rankings

Browse a broad leaderboard that spans text, web development, vision, document, image, search, video, and video-edit tasks.

Agent performance leaderboard

Inspect separate agent rankings that emphasize tool use, task completion, reliability, and steerability for real-world agentic work.

Score-based comparison view

Review ranked model lists with score displays and uncertainty ranges where available, giving more context than a simple ordered list.

Chat history search

Search chat history from the Arena interface to find earlier conversations and past interactions.

Common use cases

  • Compare model outputs

    Test how different frontier models answer the same prompt, then vote on the result to help the public ranking reflect real comparisons.

  • Review task-specific leaderboards

    Check the text, vision, web development, image, search, or video arenas when you want a quick view of which models lead in a specific task family.

  • Evaluate agents and tool use

    Use the agent leaderboard to assess how well models orchestrate tools and complete real-world agentic tasks.

  • Find past chats

    Search earlier conversations in Arena when you need to revisit prior prompts or outputs from your own chat history.

Pros and Cons

Pros

  • Covers many task areas instead of limiting comparison to one benchmark.
  • Uses public evaluation signals and voting to shape the leaderboard.
  • Provides separate views for general models and agent-focused performance.
  • Surfaces ranked lists with contextual score information rather than only a single winner.
  • Includes chat history search for revisiting earlier conversations.

Cons

  • The provided pricing URL currently returns a 404, so pricing and plan details are not available from the source.
  • The sources do not confirm integrations, API access, team features, or other platform capabilities beyond model comparison and chat history search.
  • Because the rankings are community-facing and real-world oriented, they are useful for comparison but do not replace task-specific testing in a reader’s own workflow.

FAQ

What is Arena used for?

Arena is a public leaderboard and comparison site for AI models. Users can chat with models, compare responses, vote, and review rankings across areas such as text, vision, web development, image, video, search, and agent tasks.

How does Arena compare AI models?

The site is centered on direct model comparison and public evaluation rather than a traditional software suite. The homepage highlights Battle Mode, and the leaderboard pages organize models by task area and rank them from real-world evaluation signals.

Does Arena have published pricing?

The source does not show a working pricing page. The /pricing URL currently returns a 404 page, so no plan, trial, or billing details are confirmed from the provided evidence.

Can users search past chats?

The site includes a search page for chat history, which suggests users can search past conversations. Beyond that, the provided sources do not confirm account, team, API, or workspace features.

Quick Facts

Category
AI leaderboard
Primary use
Compare and vote on AI model responses
Primary users
People evaluating LLMs, image models, code models, and agents
Source domain
lmarena.ai
Pricing
No published pricing confirmed; /pricing returns 404