{"ok":true,"entity":{"id":"ai-benchmark","name":"AI Benchmark","entityType":"concept","officialName":"AI Benchmark","canonicalName":"AI Benchmark","displayName":"AI Benchmark","category":"AI概念（モデル性能評価の標準化された基準）","shortDescription":"AIモデルの性能（言語理解・推論・コーディング等）を標準化された問題セットで定量的に評価する仕組み。MMLU等が代表例で、AI企業・ラボが自社モデルの性能を主張する際の共通の比較基準として用いられる。","primaryCluster":"ai-concepts","verificationStatus":"draft","website":null,"updatedAt":"2026-07-22T13:12:07.131Z","secondaryClusters":[],"alias":[],"searchKeywords":["AI Benchmark","AIベンチマーク","MMLU"]},"referenceIndex":["P-01-001","P-02-001","P-04-001","P-02-002","P-04-002","P-05-001"]}