Series B

Arena, the A.I. Leaderboard, Is Now Worth $3.1 Billion

LMArena. © Screenshot / Canva
LMArena. © Screenshot / Canva

Set Trending Topics as a preferred source on Google.

Whoever grades A.I. models is apparently getting a high grade too: Arena, the platform formerly known as Chatbot Arena and LMArena, has just closed a $200 million (about €178 million) Series B at a valuation of $3.1 billion (about €2.8 billion). The round was led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z and Felicis also participating.

A Valuation That Nearly Doubled

Only in January, Arena raised $150 million (about €134 million) at a $1.7 billion valuation. At the time, its annualized revenue stood at $30 million; by June, Arena reported $100 million (about €89 million). Most of the money comes from AI Evaluations, a paid service that gives A.I. labs and companies analytics based on community ratings.

Arena began in 2023 as a research project at the University of California, Berkeley. The principle is simple: Users ask a question, two anonymous A.I. models answer, and users vote on which response is better. Millions of these duels produce a leaderboard that, according to Arena, draws tens of millions of visitors a month. “We thought it was going to be a paper, not a company,” Anastasios Angelopoulos, the chief executive, told the industry publication RuntimeWire.

A New Ranking for A.I. Agents

With the new money, Arena is building ratings for A.I. agents. A new alignment index, based on about 90,000 real sessions, measures how often agents take actions without permission, misattribute sources or claim to have completed tasks they never did. In the preliminary ranking, OpenAI models lead, with GPT-6.1 Sol on top at 87.9 points. Anthropic’s Claude Opus 5.5 places sixth. “Static benchmarks break down once models recognize they’re being tested,” Arena said in explaining its approach.

For the industry, these rankings have long been a marketing tool. New models such as Mistral Large 4 are routinely promoted with their Arena placement, and benchmarks also play a central role in the price war between OpenAI and Anthropic.

Questions About Neutrality

That very role makes Arena a target. In the study “The Leaderboard Illusion,” researchers led by Sara Hooker of Cohere Labs accused the platform of favoring large providers. They said Meta had privately tested 27 variants before the release of Llama 4, that big labs could withdraw weak results and that they received more data than open-source providers. Ms. Hooker called it a “crisis” in A.I. evaluation. Arena denied treating anyone unfairly, said the study contained factual errors and tightened its rules after the Meta episode.

Rank My Startup: Erobere die Liga der Top Founder!
Advertisement
Advertisement

Specials from our Partners

Top Posts from our Network

Deep Dives

© Wiener Börse

IPO Spotlight

powered by Wiener Börse

Europe's Top Unicorn Investments 2023

The full list of companies that reached a valuation of € 1B+ this year
© Behnam Norouzi on Unsplash

Crypto Investment Tracker 2022

The biggest deals in the industry, ranked by Trending Topics
ThisisEngineering RAEng on Unsplash

Technology explained

Powered by PwC
© addendum

Inside the Blockchain

Die revolutionäre Technologie von Experten erklärt

Trending Topics Tech Talk

Der Podcast mit smarten Köpfen für smarte Köpfe
© Shannon Rowies on Unsplash

We ❤️ Founders

Die spannendsten Persönlichkeiten der Startup-Szene
Tokio bei Nacht und Regen. © Unsplash

🤖Big in Japan🤖

Startups - Robots - Entrepreneurs - Tech - Trends

Continue Reading

Newsletter

Founders Dispatch

Zwei Mal pro Woche kostenlos in die Inbox: die wichtigsten Startups, Deals und Tech-Entwicklungen aus Europa, handgeschrieben von der Redaktion.

Jederzeit abbestellbar. Mehr über den Newsletter