AIWiki
Malaysia
Back to all articles
Tools & PlatformsLLM evaluationleaderboardbenchmarking

LMArena

4 min readUpdated August 2026
LMArena
Type
Crowdsourced AI evaluation platform
Launched
May 2023 (as Chatbot Arena)
Developer
LMSYS Org, UC Berkeley; later Arena Intelligence Inc.
Method
Blind pairwise battles with Elo ratings
Rebranded
Arena (January 2026)
Related
AI benchmarking, LLM-as-a-judge

LMArena (formerly Chatbot Arena, rebranded Arena in January 2026) is a crowdsourced web platform for evaluating large language models through anonymous head-to-head comparisons. Users submit a prompt, receive responses from two unidentified models, and vote for the better answer; the votes feed an Elo-style rating system that produces a public leaderboard widely treated as a reference for real-world AI model quality.[1][2]

History and Background

LMArena was launched in May 2023 as Chatbot Arena, an academic project of the Large Model Systems Organization (LMSYS Org), a research group at the University of California, Berkeley's Sky Computing Lab. It was created by researchers including Wei-Lin Chiang and Anastasios Angelopoulos as a way to evaluate conversational AI using human preferences rather than synthetic benchmarks, and quickly grew into one of the most widely used LLM comparison platforms.[2][3]

The methodology was documented in the paper "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference" (Chiang et al., arXiv:2403.04132, March 2024).[2] In September 2024 the project moved to its own domain, lmarena.ai. In April 2025 it was incorporated as the independent company Arena Intelligence Inc., and in May 2025 it raised a US$100 million seed round at a US$600 million valuation, led by Andreessen Horowitz (a16z) and UC Investments, with participation from Lightspeed, Laude Ventures, Felicis, Kleiner Perkins and The House Fund.[3][4]

In January 2026 the company raised a US$150 million Series A round at a post-money valuation of US$1.7 billion, led by Felicis and UC Investments, and on 28 January 2026 the platform was renamed Arena, operating at arena.ai.[5]

Technology and Methodology

The core of LMArena is the blind battle: a user chats with two anonymous models without knowing which system produced which response, then votes for the winner. Results are aggregated using an Elo rating system adapted from chess, in which ratings are updated based on the expected versus actual outcome of each comparison. The platform publishes leaderboards for overall performance as well as specialized categories such as coding, mathematics, image generation and long-context tasks, and it has expanded into multimodal evaluation since mid-2024. Specialized evaluation modes include Arena-Hard, WebDev Arena and RepoChat Arena, and the underlying voting data and leaderboard rankings are released openly for research.[1][2][4]

Applications and Impact

LMArena has become a de facto reference point for comparing frontier AI models, and major AI laboratories use it to test and promote their systems. Models from providers including Google, OpenAI, Meta, Anthropic, DeepSeek and xAI have topped or climbed its leaderboard at various times, and open-weight models have competed directly with proprietary systems. The platform's results are widely quoted in product announcements and academic research, though it has also faced criticism from researchers who argue that AI labs may attempt to influence or game the rankings, an accusation the company has denied. With more than three million votes cast by May 2025 and hundreds of models evaluated, LMArena is among the largest human-preference evaluation efforts in AI.[3][4]

See Also

๐Ÿ‡ฒ๐Ÿ‡พMalaysian Context

LMArena is widely used by Malaysian AI practitioners, developers and students to compare models such as GPT-5, Claude, Gemini and DeepSeek before choosing a model for local applications, since its human-preference rankings complement traditional benchmark scores when weighing quality against cost. The platform's open leaderboard data is also referenced in Malaysian AI communities and training programmes, and Malaysian developers of open-weight Malay-language models use community evaluations to gauge how their systems compare with international models. For organisations, leaderboard results feed into model-selection decisions that must also satisfy Malaysian regulatory requirements, including the Personal Data Protection Act (PDPA) when customer data is processed.[6][7]

References

  1. โ†‘LMArena (Arena). Official leaderboard and platform. https://lmarena.ai
  2. โ†‘Chiang, W.-L., et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132. https://arxiv.org/abs/2403.04132
  3. โ†‘TechCrunch. (2025, May 21). LM Arena, the organization behind popular AI leaderboards, lands $100M. https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m
  4. โ†‘PR Newswire. (2025, May 21). LMArena Secures $100M in Seed Funding to Bring Scientific Rigor to AI Reliability. https://www.prnewswire.com/news-releases/lmarena-secures-100m-in-seed-funding-to-bring-scientific-rigor-to-ai-reliability-302462025.html
  5. โ†‘TechCrunch. (2026, January 6). LMArena lands $1.7B valuation four months after launching its product. https://techcrunch.com/2026/01/06/lmarena-lands-1-7b-valuation-four-months-after-launching-its-product
  6. โ†‘AIWiki Malaysia. AI Benchmarking. https://aiwiki.com.my/wiki/ai-benchmarking
  7. โ†‘AIWiki Malaysia. PDPA AI Compliance. https://aiwiki.com.my/wiki/pdpa-ai-compliance