LMSYS Org公式のブログ記事(https://www.lmsys.org/blog/2023-05-03-arena/)は、AI Benchmarkの差別化要因に関する第三者報道である。None of AI Benchmark's 3 existing references mention Chatbot Arena, human-preference evaluation, or Elo ratings; all 3 existing references cite only the MMLU paper.具体的には「Users chat with two anonymous models side-by-side and vote for which one is better, using the Elo rating system, a widely-used rating system in chess and other competitive games; this contrasts with traditional benchmarks like HELM and lm-evaluation-harness, which rely on academic datasets and programmatic evaluation.」といった記載がある。ただし、この差別化要因の記述は当該Sourceの時点の位置づけであり、競合状況の変化を反映しない可能性があります。 This source describes Chatbot Arena as it existed at its 2023 launch; the platform (later evolving into LMArena) has continued to change since, so specifics like vote counts or model-pool size may not reflect the current 2026 state.2026年08月27日に同Sourceを取得し、この内容を確認した。検証待ちAI Benchmark差別化要因に関する公開情報
arXiv (independent researchers)公式の学術論文(https://arxiv.org/abs/2311.09783)は、AI Benchmarkの比較優位性に関する第三者報道である。Substantively different from the Chatbot Arena finding above (data contamination/reliability concern vs. alternative evaluation methodology) and from the existing 3 references, which present MMLU/benchmarks neutrally as valid comparison tools without discussing contamination risk.具体的には「ChatGPT and GPT-4 demonstrated an exact match rate of 52% and 57%, respectively, in guessing the missing options in benchmark test data, via the paper's TS-Guessing protocol applied to benchmarks including MMLU.」といった記載がある。ただし、この比較は当該Sourceの時点の記述であり、比較対象の状況変化を反映しない可能性があります。 This paper's findings are specific to the models and benchmarks tested by its TS-Guessing protocol at time of publication (2023); contamination rates could differ for newer models or benchmarks not covered by this specific study.2026年08月27日に同Sourceを取得し、この内容を確認した。検証待ちAI Benchmark比較優位性に関する公開情報