LMSYS Chatbot Arena
lmarena.ai
A live human preference leaderboard where anonymous models compete head-to-head in blind evaluations, producing the most credible ranking of conversational AI quality available.
Why it is useful
Elo rankings based on millions of human blind comparisons are harder to game than benchmark suites, which makes this the most trustworthy signal for which models people actually prefer in practice. Useful both for choosing models and for understanding the gap between self-reported benchmark performance and real-world user preference.