Side-by-side model comparison
Users can chat with models and compare responses side by side, then vote for the answer they think is better. This turns everyday prompts into a public evaluation workflow.
Arena is a public AI model comparison and ranking platform for chatting, voting on outputs, and exploring task-specific leaderboards across text, vision, code, search, video, and agent tasks.
Arena is a public AI model comparison and ranking platform. It lets people chat with frontier models, compare responses, vote on the better answer, and help shape community leaderboards for large language models, image models, code models, and other task-specific arenas.
The site presents itself as an official leaderboard for benchmarking models across real-world evaluation settings. Dedicated leaderboard pages cover text, vision, document understanding, image generation and editing, web development, search, text-to-video, image-to-video, video editing, and agent tasks.
Users can chat with models and compare responses side by side, then vote for the answer they think is better. This turns everyday prompts into a public evaluation workflow.
Arena maintains leaderboard views across multiple arenas, including text, vision, document, image generation, image editing, web development, search, and video tasks. Each view shows ranked models and score spreads.
The Agent Arena page ranks models on real-world agentic tasks using signals such as tool reliability, task completion, and steerability. That makes it useful for judging orchestration, not just conversational quality.
The site includes a searchable chat history area for past conversations, with tabs for Agent, Battles, Search, Code, Image, Video, and Archived chats. This supports revisiting prior evaluations and examples.
Leaderboard pages expose high-level snapshots and links to deeper views, so users can start with a quick scan and drill into a specific arena when needed.
Compare responses from multiple frontier models to decide which output is most useful for a specific prompt. This is the core workflow for people evaluating general-purpose chat quality.
Review task-specific rankings when selecting a model for vision, document understanding, search, web development, or video generation work. The dedicated arenas make it easier to compare like with like.
Use the Agent Arena leaderboard to inspect how models perform when they need to orchestrate tools and complete multi-step work. The ranking emphasizes reliability, completion, and steerability.
Search past conversations and battles to revisit examples, compare previous prompts, or keep track of prior evaluations across different categories.
Arena lets you start a conversation, compare model responses, and vote on which answer is better. The site also provides search across saved chats and specialized leaderboards for agent, text, image, and other model categories.
The public site shows leaderboard pages and chat interfaces, but the pricing page at `/pricing` currently returns a 404. Based on the available sources, pricing details are not published there.
The rendered homepage warns that inputs are processed by third-party AI and that conversations and certain personal information may be disclosed to AI providers and may otherwise be disclosed publicly. Users are told not to submit sensitive information they would not want shared publicly.
The leaderboard pages are organized by task type, including text, web development, vision, document, text-to-image, image edit, image-to-webdev, search, text-to-video, image-to-video, and video edit. The agent leaderboard highlights performance for tool use, task completion, reliability, and steerability.
书生 is a general-purpose large model suite for language, multimodal, weather, ocean, industrial design, 3D, finance, and scientific discovery tasks.
OpenAI offers ChatGPT, API, Platform tools, and Codex for conversational AI, model building, and research and product updates.
小艺 is Huawei’s AI smart assistant for Q&A, writing, document reading, code help, image recognition, and file drag-and-drop.
AI Magicx is a unified AI workspace for chat, image, video, voice, music, email and developer tasks, helping teams and creators manage multiple models in one place.
Hypertype is an AI agent for industrial support teams, reducing repetitive work, integrating with existing tools, and using selected company data through consultative enterprise onboarding.
Bible.ai is a Christian AI app for scripture-based conversations, faith questions, and biblical guidance through voice or text. It is positioned for believers and the Christian community rather than general-purpose AI use.