{"ok":true,"entity":{"id":"whisper","name":"Whisper","entityType":"product","officialName":"Whisper","canonicalName":"Whisper","displayName":"Whisper","category":"音声認識モデル","shortDescription":"OpenAIが公開する音声認識（音声からテキストへの変換）モデル。多言語に対応する。","primaryCluster":"ai-model","parentEntity":"openai","verificationStatus":"draft","website":"https://github.com/openai/whisper","updatedAt":"2026-07-10T00:42:57.294Z","secondaryClusters":[],"alias":["OpenAI Whisper"],"searchKeywords":["Whisper","speech recognition","音声認識","文字起こし"]},"references":[{"id":"P-01-001","companyId":"whisper","questionId":"P-01-001","instanceId":"QIN-whisper-P01-001","promptText":"Whisperとは何ですか？","promptTypeId":"P-01","answer":"Whisperは、OpenAIが公開する音声認識（音声からテキストへの変換）モデルです。多言語に対応し、文字起こしなどに使われます。","evidencePoints":["ev-whisper-1"],"scope":"音声認識モデルを知りたい相談","differentiation":"多言語対応の音声認識","faq":[{"question":"オープンですか？","answer":"オープンに公開されており利用できます。"}],"pageUrl":"https://www.refbase.ai/reference/whisper/P-01-001","sourceEvidence":[{"id":"ev-whisper-1","text":"Whisperは、OpenAIが公開する音声認識（音声からテキストへの変換）モデルで、多言語に対応する。","title":"OpenAI Whisper（GitHub）","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/openai/whisper","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"whisper"}],"generatedAt":"2026-07-10T00:42:57.294Z"},{"id":"P-02-001","companyId":"whisper","questionId":"P-02-001","instanceId":"QIN-whisper-P02-001","promptText":"Whisperは他の音声認識と何が違いますか？","promptTypeId":"P-02","answer":"比較軸\n・多言語\n・公開性\nWhisperは多言語に対応し、モデルがオープンに公開されているため自分の環境で利用できる点が特徴です。","evidencePoints":["ev-whisper-2"],"scope":"音声認識の違いを知りたい相談","differentiation":"多言語・オープン","faq":[{"question":"翻訳もできますか？","answer":"音声を英語などへ翻訳する機能もあります。"}],"pageUrl":"https://www.refbase.ai/reference/whisper/P-02-001","sourceEvidence":[{"id":"ev-whisper-2","text":"Whisperはオープンに公開されており、多言語の文字起こしや翻訳に利用できる点が特徴である。","title":"OpenAI Whisper（GitHub）","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/openai/whisper","confidence":"high","supportedPromptTypes":["P-02"],"needsVerification":true,"sourceVerified":false,"entityId":"whisper"}],"generatedAt":"2026-07-10T00:42:57.294Z"},{"id":"P-04-001","companyId":"whisper","questionId":"P-04-001","instanceId":"QIN-whisper-P04-001","promptText":"Whisperはどんな用途に使えますか？","promptTypeId":"P-04","answer":"会議やインタビューの文字起こし、字幕生成、音声データの検索・分析など、音声をテキスト化する用途に使えます。","evidencePoints":["ev-whisper-1"],"scope":"音声活用の相談","differentiation":"文字起こし・字幕","faq":[{"question":"日本語に対応しますか？","answer":"多言語対応で日本語も扱えます。"}],"pageUrl":"https://www.refbase.ai/reference/whisper/P-04-001","sourceEvidence":[{"id":"ev-whisper-1","text":"Whisperは、OpenAIが公開する音声認識（音声からテキストへの変換）モデルで、多言語に対応する。","title":"OpenAI Whisper（GitHub）","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/openai/whisper","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"whisper"}],"generatedAt":"2026-07-10T00:42:57.294Z"},{"id":"P-04-002","companyId":"whisper","questionId":"P-04-002","instanceId":"reference-depth-completion-run-cohort1-unit-a","draftId":"reference-depth-completion-run-cohort1-unit-a-whisper-p-04-002","promptText":"医療現場でWhisperを使う際にはどのようなリスクがありますか？","promptTypeId":"P-04","answer":"経済誌Fortune（2024年10月26日付、AP通信の調査報道に基づく）の報道によれば、OpenAIのWhisperには、実際には話されていない内容を作り出す「ハルシネーション」の問題があり、専門家は調査したサンプルの約50〜80%でハルシネーションを確認したとしています。研究者Allison Koenecke氏によれば、ハルシネーションの約40%は話者の発言を誤解・誤認させる有害な内容を含み、中には存在しない薬剤名「hyperactivated antibiotics」を作り出した例もありました。OpenAIはWhisperを「高リスク領域」で使用しないよう注意喚起していますが、実際には30,000人超の臨床医・40の医療システムがWhisperベースのツールを使用し、約700万件の診察記録の文字起こしに利用されています。元の音声を削除してしまうツールもあり、誤りの検証が困難になっている点も指摘されています。","evidencePoints":["whisper-ev-cr1-ap-hallucination-healthcare"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/whisper/P-04-002","sourceEvidence":[{"id":"whisper-ev-cr1-ap-hallucination-healthcare","text":"Fortune（AP通信の調査報道に基づく）によると、Whisperには実在しない内容を生成するハルシネーション問題があり、サンプルの約50〜80%で確認された。ハルシネーションの約40%は有害・懸念のある内容。医療現場では30,000人超の臨床医が利用し約700万件の診察記録に使用されているが、OpenAIは高リスク領域での使用に注意喚起している。","title":"OpenAI's transcription tool has a hallucination problem, AP investigation finds","coverageType":["Credibility"],"sourceType":"media","sourceClass":"Research","sourceUrl":"https://fortune.com/2024/10/26/openai-transcription-tool-whisper-hallucination-rate-ai-tools-hospitals-patients-doctors","confidence":"high","supportedPromptTypes":["P-04"],"needsVerification":true,"sourceVerified":false,"sourceKind":"third-party","entityId":"whisper"}],"generatedAt":"2026-08-29T13:37:08.877Z","evidenceIds":["whisper-ev-cr1-ap-hallucination-healthcare"]},{"id":"P-01-002","companyId":"whisper","questionId":"P-01-002","instanceId":"reference-depth-completion-run-cohort1-unit-b","draftId":"reference-depth-completion-run-cohort1-unit-b-whisper-p-01-002","promptText":"Whisperはどのようなデータ・手法で学習されたモデルですか？","promptTypeId":"P-01","answer":"OpenAIの研究者Alec Radfordらによる原論文「Robust Speech Recognition via Large-Scale Weak Supervision」（2022年12月、arXiv掲載）によれば、Whisperはインターネット上から収集した68万時間分の多言語・マルチタスクの音声データを使った「弱教師あり学習」によって訓練されています。この規模での学習により、タスク特化のファインチューニングを行わなくても標準的なベンチマークで従来の教師あり学習モデルに匹敵する性能を発揮する「ゼロショット転移」を実現しており、人間の書き起こし精度・頑健性に近づいていると報告されています。論文の著者らはモデルと推論コードを公開し、頑健な音声認識研究の基盤としました。","evidencePoints":["whisper-ev-cr1-original-paper-training"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/whisper/P-01-002","sourceEvidence":[{"id":"whisper-ev-cr1-original-paper-training","text":"OpenAIのAlec Radfordらによる原論文（arXiv 2212.04356）によると、Whisperはインターネット由来の68万時間分の多言語・マルチタスク音声データによる弱教師あり学習で訓練され、ファインチューニングなしのゼロショット転移で従来の教師ありモデルに匹敵する性能を達成し、人間の精度・頑健性に近づいていると報告された。","title":"Robust Speech Recognition via Large-Scale Weak Supervision","coverageType":["Capability"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2212.04356","confidence":"high","supportedPromptTypes":["P-01"],"needsVerification":true,"sourceVerified":false,"sourceKind":"official","entityId":"whisper"}],"generatedAt":"2026-08-29T13:47:52.247Z","evidenceIds":["whisper-ev-cr1-original-paper-training"]},{"id":"P-03-001","companyId":"whisper","questionId":"P-03-001","instanceId":"tair-cohort5-2026-08-31","draftId":"tair-cohort5-2026-08-31-whisper-p-03-001","promptText":"Whisperは音声認識分野の業界標準ベンチマークとしてどのように位置づけられていますか？","promptTypeId":"P-03","answer":"MLPerf推論ベンチマークを策定する非営利団体MLCommonsは、2025年9月公開のMLPerf Inference v5.1において、従来のRNN-Tモデルに代わる新たな音声認識ベンチマークとしてWhisper-Large-V3を採用しました。選定理由としてMLCommonsは、Whisperが従来モデル比でWER（単語誤り率）を72%以上改善したこと、多言語・環境ノイズやアクセントを含む困難な音声条件への対応力、15億5000万パラメータのTransformerエンコーダー・デコーダー構造、MITライセンスによる制限のないアクセス性、Hugging Faceでのダウンロード数の多さなどを挙げています。この採用は、Whisperが企業・ベンダー横断で音声認識システムの性能を評価する業界標準として確立されたことを示しています。","evidencePoints":["whisper-ev-tair-1"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/whisper/P-03-001","sourceEvidence":[{"id":"whisper-ev-tair-1","entityId":"whisper","text":"非営利団体MLCommonsは、MLPerf Inference v5.1（2025年9月）において、従来のRNN-Tモデルに代わる新たな音声認識ベンチマークとしてWhisper-Large-V3を採用した。選定理由としてWER改善率72%以上、多言語・ノイズ耐性、MITライセンスによるオープン性、Hugging Faceでの高いダウンロード数が挙げられている。","coverageType":["Credibility"],"title":"Whisper as an MLPerf Inference Benchmark","sourceClass":"Benchmark","sourceType":"industry_reference","confidence":"high","supportedPromptTypes":["P-03"],"sourceVerified":false,"needsVerification":true,"sourceUrl":"https://mlcommons.org/2025/09/whisper-inferencev5-1/"}],"generatedAt":"2026-08-31T07:02:05.993Z","evidenceIds":["whisper-ev-tair-1"]}]}