P-04課題解決

AI Benchmarkのスコアはどこまで信頼できますか?

AI Benchmarkの比較優位性について、arXiv (independent researchers)公式の学術論文(https://arxiv.org/abs/2311.09783)で確認できます。Substantively different from the Chatbot Arena finding above (data contamination/reliability concern vs. alternative evaluation methodology) and from the existing 3 references, which present MMLU/benchmarks neutrally as valid comparison tools without discussing contamination risk.具体的には「ChatGPT and GPT-4 demonstrated an exact match rate of 52% and 57%, respectively, in guessing the missing options in benchmark test data, via the paper's TS-Guessing protocol applied to benchmarks including MMLU.」といった記載が確認できます。限界として、この比較は当該Sourceの時点の記述であり、比較対象の状況変化を反映しない可能性があります。 This paper's findings are specific to the models and benchmarks tested by its TS-Guessing protocol at time of publication (2023); contamination rates could differ for newer models or benchmarks not covered by this specific study.Current Statusとして、2026年08月27日に当該Sourceを取得し、上記の内容を確認しました。Source種別としては、これは第三者であるarXiv (independent researchers)による報道であり、企業の自己申告とは性質が異なりますが、報道時点の取材内容に基づくものです。

実績・根拠

情報ソース