P-02比較・違い

AI Benchmarkには、MMLUのような固定テストセット型以外にどのような評価方式がありますか?

AI Benchmarkの差別化要因について、LMSYS Org公式のブログ記事(https://www.lmsys.org/blog/2023-05-03-arena/)で確認できます。None of AI Benchmark's 3 existing references mention Chatbot Arena, human-preference evaluation, or Elo ratings; all 3 existing references cite only the MMLU paper.具体的には「Users chat with two anonymous models side-by-side and vote for which one is better, using the Elo rating system, a widely-used rating system in chess and other competitive games; this contrasts with traditional benchmarks like HELM and lm-evaluation-harness, which rely on academic datasets and programmatic evaluation.」といった記載が確認できます。限界として、この差別化要因の記述は当該Sourceの時点の位置づけであり、競合状況の変化を反映しない可能性があります。 This source describes Chatbot Arena as it existed at its 2023 launch; the platform (later evolving into LMArena) has continued to change since, so specifics like vote counts or model-pool size may not reflect the current 2026 state.Current Statusとして、2026年08月27日に当該Sourceを取得し、上記の内容を確認しました。Source種別としては、これは第三者であるLMSYS Orgによる報道であり、企業の自己申告とは性質が異なりますが、報道時点の取材内容に基づくものです。

実績・根拠

情報ソース