Arena Free

-

Arena (formerly LMArena / LMSYS Chatbot Arena) is a community-driven AI review platform launched by the UC Berkeley research team. Through the Battle mode and community voting mechanism, users can compare the output quality of different models in real scenarios, forming a public ranking based on 82 million+ real human preferences. Covering multi-modal evaluation dimensions such as text, code, images, and video agents, it also provides commercial evaluation services for enterprises and model laboratories.

Arena Product Interface

Arena

Core parameters and statistics of Arena

Arena (formerly LMArena / LMSYS Chatbot Arena) is a community-driven AI evaluation platform founded by the UC Berkeley research team. It is officially positioned as a "community-powered platform for understanding AI performance in the real world". It is neither an AI dialogue assistant nor a model training framework, but an evaluation infrastructure that connects "model producers" and "model consumers" - through 82 million+ real human votes, it forms a public ranking covering multi-modal dimensions such as text, code, images, and video agents.

Projects Public Information
Official positioning Community-powered AI evaluation platform
Predecessor LMSYS Chatbot Arena → LMArena (2025) → Arena (2026)
Evaluation dimensions Text, code (WebDev/Image-to-WebDev), image generation/editing, video generation/editing Vision, Document, Search, Agent
Community size 10 million + monthly active visitors 700 million + conversations 82 million + votes
Business Model Free Community Tier + Enterprise Evaluation Service (AI Evaluations)
Financing situation $100M Seed (2025.05) → $150M Series A (2026.01)
Annualized revenue $100M ARR (as of 2026.06, enterprise services have been online for 8 months)
Place of Belonging US (Berkeley, California)
Support Platform Web

A brief comment: Arena is essentially an "AI model evaluation field" - it does not allow users to choose models and chat, but through Battle mode and community voting, it changes "which model is better" from a subjective impression into a quantifiable community consensus.

Core mechanism: After the user inputs the prompt word, the system anonymously displays the output of the two models, and the user votes for the better answer based on his or her own judgment. Voting results drive public rankings, while open source data sets are available for academic and industrial research. This mechanism makes the evaluation results closer to real human preferences rather than the overfitting indicators of static benchmarks.

Revenue Structure: Arena will achieve $100M annualized revenue in June 2026, only 8 months after the enterprise evaluation service is launched. Growth comes primarily from Model Lab’s custom assessment services and selection support for enterprise customers – the $150M Series A was co-led by Felicis and UC Investments, with participation from a16z, Kleiner Perkins, Lightspeed and others.

Arena’s users and market recognition

Arena’s market recognition comes from two clear paths: data scale effect on the community side and commercial verification on the enterprise side.

Data density on the community side: According to official disclosures, the platform has a total of 700 million + 82 million conversations + 10 million + monthly active visitors. The significance of these numbers lies not in the traffic itself, but in the scarcity of "human preference data" - the vast majority of model reviews rely on static benchmarks or LLM-as-a-judge, while Arena retains the judgment of real users on open tasks. Arena has open sourced the largest human preference dataset on HuggingFace (lmarena-ai), which is cited by multiple modeling laboratories and academic institutions.

Growth curve on the enterprise side: The enterprise evaluation service reached $100M ARR 8 months after it was launched, indicating that model laboratories and large enterprises have a rigid demand for "independent, transparent evaluation based on real human feedback". 400+ new models have been publicly evaluated through Arena, covering both open source and closed source models. Agent Mode has reached 5 million+ monthly interaction rounds within one month of its launch, with a week-on-week increase of 10%, which proves that multi-step Agent evaluation is one of the most urgent gaps in the current market.

Academic influence: The Arena team has published evaluation methodology papers at top conferences such as ICML 2024 (Chatbot Arena paper), NeurIPS 2023 (LLM-as-a-judge), ICLR 2025 (RouteLLM), ICML 2025 (Prompt-to-Leaderboard, Arena-Hard), and its ranking method has become one of the de facto standards in the community.

Prerequisites: The credibility of the evaluation platform is based on the diversity of voting samples, the neutrality of prompt words, and the statistical rigor of the ranking method. Platform users are mainly technical people (the X/Discord community is active), and the sample bias may be biased towards specific usage scenarios. This is a factor that needs to be considered when interpreting the rankings.

Arena’s cost advantage

  • C-side/Individual: Usually a free version is provided to experience the core functions, and high-frequency use requires a paid package subscription.
  • API/Developer: Billed by call volume, suitable for development teams that can be flexibly integrated into their own systems.
  • Enterprise/Privatized: Contact the business owner for customized quotation and deployment plan. The specific price is subject to the official real-time pricing page.

Main features of Arena

Arena's capability is not "a chat box plus a ranking list", but a comprehensive system built around "evaluation-routing-ranking-data feedback":

  • Battle mode (anonymous model comparison): The user inputs a prompt word, two anonymous models output it at the same time, and the user votes to choose the better answer. This is the core data collection mechanism of Arena. The key design is anonymity to eliminate brand bias, random pairing to reduce ranking noise, and the ELO scoring system provides stable rankings across time. Hidden linkage: Each vote generates three values ​​at the same time - updating the rankings, training the Max routing model, and expanding the open source data set. It completes the three tasks of data collection, routing optimization and community contribution in one operation.

  • Multi-dimensional ranking system: Not just an "overall ranking", but divided into 13+ subdivided rankings according to capabilities and modes: Overall, Agent, Text, WebDev, Image-to-WebDev, Text-to-Image, Image Edit, Text-to-Video, Image-to-Video, Video Edit, Vision, Document, Search, etc. Expert View: The value of the segmentation list is to eliminate confusion - a model that ranks high in Overall may perform mediocrely in WebDev scenarios. The segmentation list changes the selection from "selecting the best model" to "selecting the model that is most suitable for this task."

  • Max Model Router: An intelligent routing engine trained on the community’s 5 million+ voting data that automatically distributes user tips to the most appropriate model. Engineering Implications: Max is essentially a "model recommendation system" based on human preference training. Its existence means that Arena is not only an evaluation platform, but also gradually becoming the decision-making layer of model scheduling.

  • Agent Mode (Multi-step Task Evaluation): A new evaluation mode launched in June 2026, measuring objective task completion rate and hallucination rate for Agent scenarios. Unlike traditional static benchmarks, Agent Mode evaluates a model's ability to perform multi-step inference and tool invocation in an open context. It has reached 5 million+ monthly interactions within one month of its launch, indicating that the market demand for Agent evaluation far exceeds expectations.

  • Enterprise Evaluation Services (AI Evaluations): Provide customized evaluations for model laboratories and enterprises, covering pre-release model verification, competitive product comparison, and fine-grained capability analysis. Core Value: Internal evaluation teams often face the problems of insufficient sample size and fixed evaluation dimensions. Arena’s enterprise service solves these two bottlenecks by accessing a real human voting pool.

Arena’s model and version evolution

Arena is not the AI model itself, so its version evolution is reflected in the continuous expansion of the platform capability layer:

  • 2025.09 Enterprise evaluation service launched: Officially entering the commercialization stage from a pure community project, launching AI Evaluations customized evaluation service.
  • 2025.11 Arena Expert: The expert-level evaluation mode is introduced, the evaluation methodology and ranking algorithm are systematically upgraded, and the ranking list is more sensitive to small differences between models.
  • 2026.01 Brand Upgrade + Series A: LMArena is renamed Arena, repositioning itself from a "chatbot arena" to a broader AI evaluation platform. Completed $150M Series A financing during the same period.
  • 2026.02 Max Model Router: Release of an intelligent routing engine trained based on 5 million+ community votes, marking the transition of Arena's capabilities from "passive evaluation" to "active scheduling".
  • 2026.05 Multimodal Extension: Added Fullstack Code Arena subdivision rankings (WebDev, Image-to-WebDev, etc.), released Multimodal Max to support multimodal model routing.
  • 2026.06 Agent Mode: Online Agent evaluation capability, supporting open multi-step task evaluation, directly responding to the industry’s urgent need for Agent reliability evaluation.
  • 2026.07 Factual Accuracy Ranking: Added Factuality Leaderboard, introducing the dimension of objective factual correctness in addition to human preferences, in response to the long-standing pain point of "difficulty in hallucination evaluation" in AI models.

Evolution Main Line: Starting from "text dialogue evaluation", it successively covers dimensions such as code, images, video agents, and practicality, while extending from "pure evaluation" to the business relationship of "evaluation + routing + enterprise services". Each version update corresponds to a point of anxiety in the industry regarding AI evaluation.

Arena’s technical advantages

The technical value of Arena is not in the performance of a single model, but in the design depth of the evaluation engineering system and the scale effect of the data flywheel:

Statistics rigor of the ELO ranking system: Arena uses an improved ELO scoring system to reduce ranking bias through anonymous pairing, random order, and sufficient sampling. Its ranking methodology paper has been peer-reviewed by ICML and is statistically more reliable than simple win rate or average rating. Causal chain: Rigorous ranking methodology → More credible rankings → The community is more willing to participate in voting → More data → More accurate rankings.

The scarcity of large-scale human preference data: 700 million+ conversations and 82 million+ votes make up one of the largest open human preference datasets in the world. The value of these data lies in "open tasks + real users + multi-dimensional preferences", which is different from the possible over-fitting risk of static benchmarks (such as MMLU, GSM8K). The open source dataset at HuggingFace has become an evaluation aid resource for several model laboratories.

Engineering reuse of Max routing engine: Max's recommendation model is essentially a by-product of community voting data - the same data stream serves both ranking updates and routing model training. This forms a positive feedback loop of "more votes → more accurate routing → better user experience → more users vote". From a cost perspective, the marginal training cost of routing capabilities is extremely low because training data is generated by free users every moment.

Unified architecture for multi-modal evaluation: Different modalities such as text, code, image, and video Agent share the same basic framework of "anonymous comparison + human voting + ELO ranking", which reduces the engineering cost of adding new evaluation dimensions. The special feature of Agent Mode is the introduction of objective completion rate measurement, which adds quantifiable task success rate in addition to subjective preferences. This design idea provides a reference for subsequent evaluation standardization.

Technical Constraints: The core bottleneck of the evaluation system is not the technical implementation, but the governance of adversarial manipulation - how to prevent the model laboratory from manipulating rankings through targeted optimization of specific prompt words, how to detect vote brushing, and how to maintain the neutrality and freshness of the prompt vocabulary. These issues will become more prominent as communities continue to grow in size.

How to use Arena

Arena is mainly accessed through the web, providing multiple entrances to suit the needs of different roles:

How to use Suitable for the crowd Operation path Typical steps
Battle mode All users arena.ai homepage → Select task type → Enter prompt word → Compare anonymous output → Vote Complete a model comparison within 1 minute
Direct conversation Users who need a specific model arena.ai → Direct Mode → Select model → Enter conversation Equivalent to a multi-model chat client
View the rankings Technical decision-makers arena.ai/leaderboard → Select dimensions (Overall/Agent/Code, etc.) View without registration
Max routing Users who want to automatically get the best answer arena.ai → Max Mode → Enter the prompt word The system automatically selects the most appropriate model
Agent Mode Evaluate multi-step tasks arena.ai/agent → Define task goals → Observe the model execution process Support open long task evaluation

Quick Start for Executives: When technical decision-makers evaluate whether Arena can be used as an evaluation tool, there are only three steps required - ① Visit arena.ai/leaderboard to see whether the structure of the segmented rankings covers the capability dimensions you care about; ② Use your own business prompt words (rather than general prompt words) to compare 10-20 times in Battle mode to feel whether the difference in results is consistent with the actual experience; ③ Contact [email protected] to learn about the customization scope and data isolation scheme of enterprise evaluation.

Data consumption path: Researchers and competitive product analysis teams can directly download the open source dataset (lmarena-ai) on HuggingFace, which contains complete anonymous voting records and model output for secondary analysis and custom ranking calculations.

Arena product pricing

The pricing model is subject to the official real-time page. Usually a freemium or subscription system is adopted. Basic functions can be used for free, while advanced functions or high-frequency use require paid subscriptions. It is recommended that users evaluate the optimal solution based on actual usage.

Arena application scenarios

Arena’s evaluation infrastructure capabilities can be embedded in a variety of decision-making links:

  • Model Selection and Procurement Decision: Before purchasing or accessing AI models, enterprises can quickly narrow down the candidates through Arena rankings and custom comparisons. Implementation Tip: The value of segmented rankings is greater than the overall rankings - for example, if the main scenario is front-end code generation, you should focus on the WebDev and Image-to-WebDev rankings instead of the Overall ranking. It is recommended to use Arena for preliminary screening first, and then use candidate models to conduct small-scale A/B testing on your own business data to make the final decision.

  • Iterative verification in Model Lab: The model development team obtains independent, third-party human preference assessments through Arena's enterprise evaluation service before release to verify the effect of version improvements. Cost reduction and efficiency improvement: Traditional models need to build their own evaluation team and annotation platform before releasing them, and it usually takes 3-6 months from establishment to output. Using Arena enterprise reviews, this cycle can be compressed to 2-4 weeks, and the cost of reviews is reduced from hundreds of thousands of dollars to the annual fee level agreed upon in the contract.

  • Agent and multi-step task reliability assessment: With the explosion of Agent products, traditional single-round evaluation can no longer cover multi-step reasoning and tool calling scenarios. Arena's Agent Mode provides open task contexts, which can be used to evaluate objective indicators such as the order completion rate of customer service agents and the construction success rate of code agents. Scenario Pain Points: The biggest difficulty in Agent evaluation is that the "gold standard" is difficult to define - for the same task, different execution paths may be correct. Arena Agent Mode alleviates this problem by introducing a hybrid assessment of human judgment + objective completion rates.

  • Education and Technology Evangelism: In teaching scenarios, Arena's Battle mode can visually display the differences in how different models process the same prompt word, helping students understand the boundaries of the model's capabilities. The community size of 10.1 million+ monthly active visitors itself also reflects its penetration in the education scene.

Applicable people for Arena

Arena's multi-layered capability design serves four distinct types of roles, with each layer having different depths of use and benefits:

  • AI product managers and technical decision-makers: People who need to do model selection. The core value of Arena is to provide third-party evaluation data independent of model manufacturers. Recommended path: First use the ranking list to narrow down the candidate scope (1 day), then use Battle mode to conduct small-scale verification with your own scene data (3-5 days), and finally contact the enterprise evaluation for in-depth evaluation (1-2 weeks). Unfit boundary: If your scene is highly vertical (such as medical imaging diagnosis, legal document review), Arena's general evaluation data may not be fine-grained enough, and you need to make customized evaluations on your own data sets.

  • Model developers and AI researchers: People who need to verify the iteration effect of model versions. Arena's enterprise reviews complement real-life human preference dimensions that are difficult for in-house review teams to cover. Open source datasets are also a valuable resource for research on ranking methodologies, preference alignment, and model evaluation techniques.

  • Enterprise Procurement and Compliance Team: The group of people who need to select an AI vendor for the organization. Arena’s segmented rankings and public methodology provide reusable selection demonstration materials. Not suitable for the boundary: The data from the evaluation platform does not constitute a complete assessment of the security of the model - if your scenario involves sensitive data or strict compliance requirements (such as HIPAA, GDPR), independent security audits and compliance reviews are still required.

  • AI Enthusiasts and Early Adopters: Learn about the differences in cutting-edge model capabilities through Battle mode, with free access to multiple models. Such users are also major contributors to the platform’s voting data.

Summary and Outlook of Arena

Arena's core competitiveness lies in the two-wheel model of "community-driven data flywheel + commercial evaluation service" - 82 million+ real human votes constitute an extremely high data barrier, and the rapid growth of enterprise evaluation services ($100M ARR in 8 months) verifies the monetization ability of this asset on the commercial side. The ability expansion path from text to agent to facts shows that Arena is evolving from a "model chat arena" to a "comprehensive evaluation infrastructure for AI capabilities."

Current limitations and uncertainties: ① The sample bias of the rankings - active users are mainly technical and early adopters, and may not represent the preferences of universal users; ② The risk of adversarial manipulation - as the commercial influence of the rankings grows, the model laboratory has the incentive to optimize the prompt words of the rankings, and the ranking method needs to continue to combat this Goodhart effect; ③ The privatized deployment capability is not disclosed - enterprises with high requirements for data sovereignty need to confirm whether Arena supports completely isolated evaluation contexts; ④ Multi-language coverage - the current evaluation data is mainly in English, and the reliability of rankings in non-English scenarios such as Chinese needs to be verified.

Procurement/Adoption Risk Assessment: As a reference tool for model comparison, Arena is extremely cost-effective – the free tier already covers the core comparison functionality. However, for enterprise-level procurement, it is recommended to proceed in three stages: ① First use the free tier to verify the consistency of ranking data and actual experience in 2-3 core scenarios (1 month); ② If more fine-grained evaluation capabilities are needed, contact the enterprise evaluation team to request customized solutions and reference customer cases; ③ Before signing a contract, focus on confirming data isolation, result attribution and exclusivity terms, and require an SLA covering evaluation response time and data availability guarantee. The final selection decision should not only rely on Arena’s evaluation data, but also need to be combined with the actual measurement results of its own business scenarios and independent security assessments.

Related tools: deepseek, chatgpt

How to use Arena

  • Web client: You can use it by visiting the official website and registering an account. Most functions do not require installation.
  • API access: Provides RESTful API, developers can obtain the API Key and integrate it into their own applications.

Version Info

  • Arena Platform (Agent Mode + Factuality Leaderboard) :Added a new Factuality Leaderboard, expanded Agent Mode evaluation capabilities to 5 million+ monthly interaction rounds, and continued to improve multi-modal evaluation coverage.
  • Agent Mode + Fullstack Code Arena :Released Agent Mode to support objective completion rate and hallucination rate assessment of multi-step complex tasks; launched Fullstack Code Arena rankings.
  • New Categories + Multimodal Max :Added Code Arena subdivision rankings (WebDev / Image-to-WebDev, etc.); released Multimodal Max multi-modal model routing capabilities.
  • Max Model Router :Released Max model routing engine that automatically routes user prompts to the most appropriate model based on 5 million+ community votes.
  • LMArena Rebrand → Arena :LMArena officially changed its name to Arena, completing brand upgrade and product expansion. During the same period, it completed Series A financing of US$150 million.
  • Arena Expert :Launched the Arena Expert expert-level evaluation mode and introduced more refined domain ability evaluation dimensions and ranking methodology updates.
  • Enterprise Evaluation Launch :Officially launched commercial evaluation service (AI Evaluations) to provide customized model evaluation for enterprises and model laboratories.

User Reviews

  • Loading reviews...