AIWiki
Malaysia
Back to all articles
AI Foundationslmarenachatbot arenabenchmarking

LMArena

4 min readUpdated September 2026
LMArena
Type
LLM evaluation platform
Created by
LMSYS Org and UC Berkeley researchers
Launched
April 2023 (as Chatbot Arena)
Method
Crowdsourced pairwise comparisons with statistical ranking
Renamed
LMArena (2024), Arena (January 2026)
Website
arena.ai
Related
AI Benchmarking, SWE-bench, LLM-as-a-Judge
LMArena — launched in April 2023 as Chatbot Arena, renamed LMArena in 2024 and rebranded as Arena in January 2026 — is a public, web-based platform that evaluates large language models through anonymous head-to-head comparisons. A user submits a prompt, receives answers from two unidentified models, and votes for the better response, after which the models are revealed. The resulting leaderboards have become among the most widely cited in the AI industry, used by developers, journalists and model makers themselves.[1][2]

History

Chatbot Arena was released on 24 April 2023 by members of LMSYS Org, a research collective, together with UC Berkeley's SkyLab. Its open-source, live approach contrasted with static academic benchmarks, and a March 2024 paper describing the platform reported that it had already amassed over 240,000 votes and that crowdsourced votes agreed closely with those of expert raters.[2]

The platform expanded steadily: image support followed in June 2024, and in September 2024 it moved to its own dedicated domain, lmarena.ai. In April 2025 the project incorporated as an independent company, and in May 2025 it raised a US$100 million seed round at a valuation of US$600 million. A further US$150 million Series A in January 2026 valued the company at about US$1.7 billion. On 28 January 2026 the platform was rebranded simply as Arena, having added video generation to its arenas earlier that month.[1][3][4][5]

Key Concepts and Technology

The platform's core is the anonymous "battle": two models answer the same prompt, the user votes without knowing which model produced which answer, and votes accumulate into leaderboards computed with statistical rating methods. Users can also compare chosen models directly, and separate arenas cover text, image understanding, image generation, video, web development, search and agentic tasks.[1]

Because voting is open to anyone, the platform captures preferences from a very large and diverse pool of users rather than a small panel of experts.[2] Its rankings have also influenced model releases: companies have previewed upcoming models anonymously on the platform, including DeepSeek, whose prototypes appeared months before its reasoning models drew wide attention, OpenAI's GPT-5 under the codename "summit", and Google's image model that became known as Nano Banana.[1] The methodology has attracted criticism as well — researchers have shown that organised voting can skew rankings, and commentators have questioned how well preference votes measure capability; the platform tightened its policies in 2025 after a controversy over a model version that differed from its public release.[1][6]

Applications and Impact

For users, LMArena offers a practical, continuously updated way to compare models on real questions rather than fixed test sets, and its leaderboards are frequently cited in press coverage and model announcements. Its rise reflects a broader shift in evaluation, in which live human preference data is treated as a complement — and sometimes a corrective — to academic benchmarks such as MMLU and SWE-bench.[1][2]

>See Also

🇲🇾Malaysian Context

For Malaysian developers, agencies and researchers choosing between models, rankings of this kind are commonly consulted, though global leaderboards are dominated by English-language evaluation. To fill that gap locally, Universiti Malaya and YTL AI Labs jointly released MalayMMLU in 2024, a benchmark of 24,213 questions spanning primary and secondary school subjects in Bahasa Melayu. In its published zero-shot results, the locally trained model ILMU 0.1 recorded a higher average score than GPT-4o, illustrating both the value and the limits of generic leaderboards for Malaysian use cases. The effort mirrors LMArena's community-driven approach at a national scale.[7]

References

  1. ↑Wikipedia contributors. (2026). Arena.ai (LLM platform). https://en.wikipedia.org/wiki/Arena.ai_(LLM_platform)
  2. ↑Chiang, W.-L. et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv. https://arxiv.org/abs/2403.04132
  3. ↑Wiggers, K. (2025). LM Arena, the organization behind popular AI leaderboards, lands $100M. TechCrunch. https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/
  4. ↑Reuters. (2026). AI startup LMArena triples its valuation to $1.7 billion in latest fundraise. https://www.reuters.com/technology/ai-startup-lmarena-triples-its-valuation-17-billion-latest-fundraise-2026-01-06/
  5. ↑Arena. (2026). About us. https://arena.ai/about
  6. ↑Robison, K. (2025). Meta got caught gaming AI benchmarks. The Verge. https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming
  7. ↑UMxYTL AI Labs. (2024). MalayMMLU: A Multitask Benchmark for the Low-Resource Malay Language. GitHub. https://github.com/UMxYTL-AI-Labs/MalayMMLU