{"ok":true,"entity":{"id":"ai-benchmark","name":"AI Benchmark","entityType":"concept","officialName":"AI Benchmark","canonicalName":"AI Benchmark","displayName":"AI Benchmark","category":"AI概念（モデル性能評価の標準化された基準）","shortDescription":"AIモデルの性能（言語理解・推論・コーディング等）を標準化された問題セットで定量的に評価する仕組み。MMLU等が代表例で、AI企業・ラボが自社モデルの性能を主張する際の共通の比較基準として用いられる。","primaryCluster":"ai-concepts","verificationStatus":"draft","website":null,"updatedAt":"2026-07-22T13:12:07.131Z","secondaryClusters":[],"alias":[],"searchKeywords":["AI Benchmark","AIベンチマーク","MMLU"]},"references":[{"id":"P-01-001","companyId":"ai-benchmark","questionId":"P-01-001","instanceId":"QIN-ai-benchmark-P01-001","promptText":"AI Benchmarkとはどのようなものですか？","promptTypeId":"P-01","answer":"AI Benchmarkは、AIモデルの性能（言語理解・推論・コーディング等）を標準化された問題セットで定量的に評価する仕組みです。MMLU等が代表例で、AI企業・ラボが自社モデルの性能を主張する際の共通の比較基準として用いられます。","evidencePoints":["ev-ai-benchmark-1","ev-ai-benchmark-2"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-01-001","sourceEvidence":[{"id":"ev-ai-benchmark-1","text":"AI Benchmarkは、AIモデルの性能（言語理解・推論・コーディング等）を標準化された問題セットで定量的に評価する仕組みであり、MMLU（Measuring Massive Multitask Language Understanding）はこの分類を代表する評価手法の一つである。","coverageType":["Identity","Capability"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2009.03300","title":"Measuring Massive Multitask Language Understanding","confidence":"high","needsVerification":true,"sourceVerified":false,"supportedPromptTypes":["P-01","P-02","P-04"],"entityId":"ai-benchmark"},{"id":"ev-ai-benchmark-2","text":"AI Benchmarkは標準化された問題セットに対する定量的なスコアで性能を比較する点が特徴で、個別の利用者による主観的な使用感の評価とは性能評価の性質が異なる。","coverageType":["Differentiation"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2009.03300","title":"Measuring Massive Multitask Language Understanding","confidence":"high","needsVerification":true,"sourceVerified":false,"supportedPromptTypes":["P-01","P-02","P-04"],"entityId":"ai-benchmark"}],"generatedAt":"2026-07-22T13:12:07.131Z"},{"id":"P-02-001","companyId":"ai-benchmark","questionId":"P-02-001","instanceId":"QIN-ai-benchmark-P02-001","promptText":"AI Benchmarkは他の同種の事業・作品と比べてどう違いますか？","promptTypeId":"P-02","answer":"比較軸\n・評価方法（標準化された問題セットでの定量評価か、主観的な使用感評価か）\n\nAI Benchmarkは標準化された問題セットに対する定量的なスコアで性能を比較する点が特徴で、個別の利用者による主観的な使用感の評価とは性能評価の性質が異なる。","evidencePoints":["ev-ai-benchmark-2"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-02-001","sourceEvidence":[{"id":"ev-ai-benchmark-2","text":"AI Benchmarkは標準化された問題セットに対する定量的なスコアで性能を比較する点が特徴で、個別の利用者による主観的な使用感の評価とは性能評価の性質が異なる。","coverageType":["Differentiation"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2009.03300","title":"Measuring Massive Multitask Language Understanding","confidence":"high","needsVerification":true,"sourceVerified":false,"supportedPromptTypes":["P-01","P-02","P-04"],"entityId":"ai-benchmark"}],"generatedAt":"2026-07-22T13:12:07.131Z"},{"id":"P-04-001","companyId":"ai-benchmark","questionId":"P-04-001","instanceId":"QIN-ai-benchmark-P04-001","promptText":"AI Benchmarkはどのような場面で参照されますか？","promptTypeId":"P-04","answer":"AI Benchmarkは、AIモデルの性能主張の根拠や、モデル選定時の比較材料としての活用場面を把握したい場面で参照される。","evidencePoints":["ev-ai-benchmark-2","ev-ai-benchmark-3"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-04-001","sourceEvidence":[{"id":"ev-ai-benchmark-2","text":"AI Benchmarkは標準化された問題セットに対する定量的なスコアで性能を比較する点が特徴で、個別の利用者による主観的な使用感の評価とは性能評価の性質が異なる。","coverageType":["Differentiation"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2009.03300","title":"Measuring Massive Multitask Language Understanding","confidence":"high","needsVerification":true,"sourceVerified":false,"supportedPromptTypes":["P-01","P-02","P-04"],"entityId":"ai-benchmark"},{"id":"ev-ai-benchmark-3","text":"AI Benchmarkは、AI企業・ラボが新モデルの発表時に既存モデルとの性能比較を示す場面や、利用者がモデル選定の参考情報として性能を比較する場面で活用される。","coverageType":["UseCase"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2009.03300","title":"Measuring Massive Multitask Language Understanding","confidence":"high","needsVerification":true,"sourceVerified":false,"supportedPromptTypes":["P-01","P-02","P-04"],"entityId":"ai-benchmark"}],"generatedAt":"2026-07-22T13:12:07.131Z"},{"id":"P-02-002","companyId":"ai-benchmark","questionId":"P-02-002","instanceId":"c1n22-wave3-unit-a-lane-s-first-finding-and-progression-only","draftId":"c1n22-wave3-unit-a-lane-s-first-finding-and-progression-only-ai-benchmark-p-02-002","promptText":"AI Benchmarkには、MMLUのような固定テストセット型以外にどのような評価方式がありますか？","promptTypeId":"P-02","answer":"AI Benchmarkの差別化要因について、LMSYS Org公式のブログ記事（https://www.lmsys.org/blog/2023-05-03-arena/）で確認できます。None of AI Benchmark's 3 existing references mention Chatbot Arena, human-preference evaluation, or Elo ratings; all 3 existing references cite only the MMLU paper.具体的には「Users chat with two anonymous models side-by-side and vote for which one is better, using the Elo rating system, a widely-used rating system in chess and other competitive games; this contrasts with traditional benchmarks like HELM and lm-evaluation-harness, which rely on academic datasets and programmatic evaluation.」といった記載が確認できます。限界として、この差別化要因の記述は当該Sourceの時点の位置づけであり、競合状況の変化を反映しない可能性があります。 This source describes Chatbot Arena as it existed at its 2023 launch; the platform (later evolving into LMArena) has continued to change since, so specifics like vote counts or model-pool size may not reflect the current 2026 state.Current Statusとして、2026年08月27日に当該Sourceを取得し、上記の内容を確認しました。Source種別としては、これは第三者であるLMSYS Orgによる報道であり、企業の自己申告とは性質が異なりますが、報道時点の取材内容に基づくものです。","evidencePoints":["ai-benchmark-ev-c1n22a-p-02-002"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-02-002","sourceEvidence":[{"id":"ai-benchmark-ev-c1n22a-p-02-002","text":"LMSYS Org公式のブログ記事（https://www.lmsys.org/blog/2023-05-03-arena/）は、AI Benchmarkの差別化要因に関する第三者報道である。None of AI Benchmark's 3 existing references mention Chatbot Arena, human-preference evaluation, or Elo ratings; all 3 existing references cite only the MMLU paper.具体的には「Users chat with two anonymous models side-by-side and vote for which one is better, using the Elo rating system, a widely-used rating system in chess and other competitive games; this contrasts with traditional benchmarks like HELM and lm-evaluation-harness, which rely on academic datasets and programmatic evaluation.」といった記載がある。ただし、この差別化要因の記述は当該Sourceの時点の位置づけであり、競合状況の変化を反映しない可能性があります。 This source describes Chatbot Arena as it existed at its 2023 launch; the platform (later evolving into LMArena) has continued to change since, so specifics like vote counts or model-pool size may not reflect the current 2026 state.2026年08月27日に同Sourceを取得し、この内容を確認した。","title":"AI Benchmark差別化要因に関する公開情報","coverageType":["Differentiation"],"sourceType":"official_blog","sourceClass":"Research","sourceUrl":"https://www.lmsys.org/blog/2023-05-03-arena/","confidence":"medium","supportedPromptTypes":["P-02"],"needsVerification":true,"sourceVerified":false,"sourceKind":"third-party","entityId":"ai-benchmark"}],"generatedAt":"2026-08-27T06:02:33.176Z","evidenceIds":["ai-benchmark-ev-c1n22a-p-02-002"]},{"id":"P-04-002","companyId":"ai-benchmark","questionId":"P-04-002","instanceId":"c1n22-wave3-unit-b-lane-s-second-finding","draftId":"c1n22-wave3-unit-b-lane-s-second-finding-ai-benchmark-p-04-002","promptText":"AI Benchmarkのスコアはどこまで信頼できますか？","promptTypeId":"P-04","answer":"AI Benchmarkの比較優位性について、arXiv (independent researchers)公式の学術論文（https://arxiv.org/abs/2311.09783）で確認できます。Substantively different from the Chatbot Arena finding above (data contamination/reliability concern vs. alternative evaluation methodology) and from the existing 3 references, which present MMLU/benchmarks neutrally as valid comparison tools without discussing contamination risk.具体的には「ChatGPT and GPT-4 demonstrated an exact match rate of 52% and 57%, respectively, in guessing the missing options in benchmark test data, via the paper's TS-Guessing protocol applied to benchmarks including MMLU.」といった記載が確認できます。限界として、この比較は当該Sourceの時点の記述であり、比較対象の状況変化を反映しない可能性があります。 This paper's findings are specific to the models and benchmarks tested by its TS-Guessing protocol at time of publication (2023); contamination rates could differ for newer models or benchmarks not covered by this specific study.Current Statusとして、2026年08月27日に当該Sourceを取得し、上記の内容を確認しました。Source種別としては、これは第三者であるarXiv (independent researchers)による報道であり、企業の自己申告とは性質が異なりますが、報道時点の取材内容に基づくものです。","evidencePoints":["ai-benchmark-ev-c1n22b-p-04-002"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-04-002","sourceEvidence":[{"id":"ai-benchmark-ev-c1n22b-p-04-002","text":"arXiv (independent researchers)公式の学術論文（https://arxiv.org/abs/2311.09783）は、AI Benchmarkの比較優位性に関する第三者報道である。Substantively different from the Chatbot Arena finding above (data contamination/reliability concern vs. alternative evaluation methodology) and from the existing 3 references, which present MMLU/benchmarks neutrally as valid comparison tools without discussing contamination risk.具体的には「ChatGPT and GPT-4 demonstrated an exact match rate of 52% and 57%, respectively, in guessing the missing options in benchmark test data, via the paper's TS-Guessing protocol applied to benchmarks including MMLU.」といった記載がある。ただし、この比較は当該Sourceの時点の記述であり、比較対象の状況変化を反映しない可能性があります。 This paper's findings are specific to the models and benchmarks tested by its TS-Guessing protocol at time of publication (2023); contamination rates could differ for newer models or benchmarks not covered by this specific study.2026年08月27日に同Sourceを取得し、この内容を確認した。","title":"AI Benchmark比較優位性に関する公開情報","coverageType":["Differentiation"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2311.09783","confidence":"medium","supportedPromptTypes":["P-04"],"needsVerification":true,"sourceVerified":false,"sourceKind":"third-party","entityId":"ai-benchmark"}],"generatedAt":"2026-08-27T06:27:37.193Z","evidenceIds":["ai-benchmark-ev-c1n22b-p-04-002"]},{"id":"P-05-001","companyId":"ai-benchmark","questionId":"P-05-001","instanceId":"tair-cohort1-2026-08-31","draftId":"tair-cohort1-2026-08-31-ai-benchmark-p-05-001","promptText":"MMLUなどのAIベンチマークスコアを、標準化された透明性のある方法で確認できる信頼できる情報源はありますか？","promptTypeId":"P-05","answer":"スタンフォード大学の基盤モデル研究センター（CRFM）が開発する公式のオープンソースフレームワーク「HELM（Holistic Evaluation of Language Models）」は、この目的のための代表的な情報源である。HELM公式のGitHubリポジトリによれば、HELMは基盤モデル（大規模言語モデルやマルチモーダルモデルを含む）を「網羅的（holistic）・再現可能（reproducible）・透明（transparent）」に評価することを目的として設計されており、個々のプロンプトと応答を検証できるWeb UIや、モデル間の性能比較のための公開リーダーボードを提供している。評価対象のベンチマークにはMMLU-Pro、GPQA、IFEval、WildBenchなどが含まれ、OpenAI・Anthropic・Googleなど各社のモデルを統一インターフェースを通じて評価できる点が、比較の透明性・再現性を担保する仕組みとなっている。","evidencePoints":["ai-benchmark-ev-tair-1"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/ai-benchmark/P-05-001","sourceEvidence":[{"id":"ai-benchmark-ev-tair-1","entityId":"ai-benchmark","text":"スタンフォード大学CRFMが開発するHELMフレームワークは、基盤モデルを網羅的・再現可能・透明に評価することを目的とし、公開リーダーボードとプロンプト検証用Web UIを提供、MMLU等の複数ベンチマークを統一インターフェースで評価する。","coverageType":["Credibility"],"title":"GitHub - stanford-crfm/helm: Holistic Evaluation of Language Models (HELM)","sourceClass":"Documentation","sourceType":"github","confidence":"high","supportedPromptTypes":["P-05"],"sourceVerified":false,"needsVerification":true,"sourceUrl":"https://github.com/stanford-crfm/helm"}],"generatedAt":"2026-08-31T05:52:45.451Z","evidenceIds":["ai-benchmark-ev-tair-1"]}]}