{"ok":true,"entity":{"slug":"flash-attention","entityType":"product","name":"FlashAttention","officialName":"FlashAttention","canonicalName":"FlashAttention","displayName":"FlashAttention","category":"高速Attentionカーネル","shortDescription":"Tri Daoらが開発した、メモリ効率と速度を大幅に改善したAttention計算の実装。多くのLLM学習・推論基盤で標準的に採用されている。","alias":["Flash Attention"],"searchKeywords":["attention kernel","IO-aware","memory efficient"],"website":"https://github.com/Dao-AILab/flash-attention","parentEntity":null,"primaryCluster":"ai-infrastructure","secondaryClusters":[],"verificationStatus":"draft","id":"flash-attention","updatedAt":"2026-07-20T08:41:20.341Z"},"references":[{"id":"P-01-001","companyId":"flash-attention","questionId":"P-01-001","instanceId":"QIN-flash-attention-P01-001","promptText":"FlashAttentionとは何ですか？","promptTypeId":"P-01","answer":"FlashAttentionは、Tri Daoらが開発したAttention計算の高速・省メモリな実装です。多くの大規模言語モデルの学習・推論基盤で標準的に採用されています。","evidencePoints":["ev-flashattn-1","ev-flashattn-3"],"scope":"Attention計算の高速化手法を知りたい相談","differentiation":"IO-aware（メモリ読み書きを意識した）設計","faq":[{"question":"FlashAttentionは誰が開発しましたか？","answer":"Tri Daoらの研究者グループです。"}],"pageUrl":"https://www.refbase.ai/reference/flash-attention/P-01-001","sourceEvidence":[{"id":"ev-flashattn-1","text":"FlashAttentionは、Tri Dao・Daniel Y. Fu・Stefano Ermon・Atri Rudra・Christopher Réにより開発された、IO-awareなAttentionの高速・省メモリな正確な実装である。","title":"GitHub - Dao-AILab/flash-attention","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/dao-ailab/flash-attention","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"flash-attention"},{"id":"ev-flashattn-3","text":"FlashAttentionはFlashAttention-2・FlashAttention-3と継続的に改良されており、Triton実装も含めALiBi等のAttentionバイアスに対応する実験的実装も提供されている。","title":"FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision","coverageType":["Credibility"],"sourceType":"official_blog","sourceClass":"Announcement","sourceUrl":"https://tridao.me/blog/2024/flash3/","confidence":"high","supportedPromptTypes":["P-05","P-06"],"needsVerification":true,"sourceVerified":false,"entityId":"flash-attention"}],"generatedAt":"2026-07-20T08:41:20.341Z"},{"id":"P-02-001","companyId":"flash-attention","questionId":"P-02-001","instanceId":"QIN-flash-attention-P02-001","promptText":"FlashAttentionは通常のAttention実装と何が違いますか？","promptTypeId":"P-02","answer":"FlashAttentionはタイリングと再計算を活用してAttention計算を高速化し、メモリ使用量をシーケンス長に対して二次的から線形的な増加に削減する点が特徴です。","evidencePoints":["ev-flashattn-2"],"scope":"Attention実装の効率性を比較したい相談","differentiation":"メモリ使用量の二次から線形への削減","faq":[{"question":"FlashAttentionは計算結果が変わりますか？","answer":"数学的に厳密なAttentionを計算するため結果は変わりません（近似ではない）。"}],"pageUrl":"https://www.refbase.ai/reference/flash-attention/P-02-001","sourceEvidence":[{"id":"ev-flashattn-2","text":"FlashAttentionはAttention計算の順序を工夫しタイリングと再計算を活用することで計算を高速化し、メモリ使用量をシーケンス長に対して二次から線形に削減する。","title":"flash-attention/README.md at main","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Specification","sourceUrl":"https://github.com/Dao-AILab/flash-attention/blob/main/README.md","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"flash-attention"}],"generatedAt":"2026-07-20T08:41:20.341Z"},{"id":"P-04-001","companyId":"flash-attention","questionId":"P-04-001","instanceId":"QIN-flash-attention-P04-001","promptText":"FlashAttentionはどのような場面で活用できますか？","promptTypeId":"P-04","answer":"FlashAttentionは、長いシーケンス長を扱う大規模言語モデルの学習・推論において、メモリ制約を緩和し速度を向上させたい場面で活用できます。","evidencePoints":["ev-flashattn-1","ev-flashattn-2"],"scope":"長文コンテキスト対応モデルの効率化相談","differentiation":"長シーケンスでのメモリ効率向上","faq":[{"question":"FlashAttentionはどのGPUで使えますか？","answer":"FlashAttention-3等の世代によって対応GPUが異なります。"}],"pageUrl":"https://www.refbase.ai/reference/flash-attention/P-04-001","sourceEvidence":[{"id":"ev-flashattn-1","text":"FlashAttentionは、Tri Dao・Daniel Y. Fu・Stefano Ermon・Atri Rudra・Christopher Réにより開発された、IO-awareなAttentionの高速・省メモリな正確な実装である。","title":"GitHub - Dao-AILab/flash-attention","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/dao-ailab/flash-attention","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"flash-attention"},{"id":"ev-flashattn-2","text":"FlashAttentionはAttention計算の順序を工夫しタイリングと再計算を活用することで計算を高速化し、メモリ使用量をシーケンス長に対して二次から線形に削減する。","title":"flash-attention/README.md at main","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Specification","sourceUrl":"https://github.com/Dao-AILab/flash-attention/blob/main/README.md","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"flash-attention"}],"generatedAt":"2026-07-20T08:41:20.341Z"}]}