{"ok":true,"entity":{"slug":"llava","entityType":"product","name":"LLaVA","officialName":"LLaVA","canonicalName":"LLaVA","displayName":"LLaVA","category":"視覚言語モデル（オープン）","shortDescription":"視覚エンコーダと言語モデルを組み合わせたオープンの視覚言語モデル。Visual Instruction Tuningの手法により、当初はLlama系の言語モデルを用いて構築された。コード・ウェイトが公開されている。","alias":["Large Language and Vision Assistant"],"searchKeywords":["マルチモーダル","視覚言語モデル","VLM","オープンモデル"],"website":"https://llava-vl.github.io","parentEntity":null,"primaryCluster":"ai-open-model","secondaryClusters":["ai-model"],"id":"llava","verificationStatus":"draft","updatedAt":"2026-07-18T11:08:14.467Z"},"references":[{"id":"P-01-001","companyId":"llava","questionId":"P-01-001","instanceId":"QIN-llava-P01-001","promptText":"LLaVAとはどのようなモデルですか？","promptTypeId":"P-01","answer":"LLaVAは視覚エンコーダと言語モデルを組み合わせたオープンの視覚言語モデルです。Visual Instruction Tuningという手法により当初はLlama系の言語モデルを用いて構築され、画像についての対話や指示応答ができます。コード・ウェイトが公開されています。","evidencePoints":["ev-llava-1","ev-llava-2"],"scope":"オープンなマルチモーダルモデルを探す相談","differentiation":"オープン構成要素の組み合わせによる視覚言語モデル","faq":[{"question":"LLaVAは何の略ですか？","answer":"Large Language and Vision Assistantの略です。"}],"pageUrl":"https://www.refbase.ai/reference/llava/P-01-001","sourceEvidence":[{"id":"ev-llava-1","text":"LLaVAは視覚エンコーダと言語モデルを接続したオープンの視覚言語モデルで、画像についての対話・指示応答ができる。","title":"LLaVA — Project Page","coverageType":["Identity","Capability"],"sourceType":"official_site","sourceClass":"Profile","sourceUrl":"https://llava-vl.github.io","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"},{"id":"ev-llava-2","text":"LLaVAのコード・学習データ・ウェイトはGitHubで公開されており、オープンな視覚言語モデルとして再現・拡張が可能である。","title":"GitHub — haotian-liu/LLaVA","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/haotian-liu/LLaVA","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"}],"generatedAt":"2026-07-18T11:08:14.467Z"},{"id":"P-02-001","companyId":"llava","questionId":"P-02-001","instanceId":"QIN-llava-P02-001","promptText":"LLaVAとクローズドなマルチモーダルAIの違いは何ですか？","promptTypeId":"P-02","answer":"クローズドなマルチモーダルAIはAPI経由でのみ利用できるのに対し、LLaVAはコード・学習データ・ウェイトが公開されており、自分の環境で実行・再現・拡張できます。研究コミュニティ発のオープンモデルとして、NeurIPS採択論文で手法が公開されている点も特徴です。","evidencePoints":["ev-llava-2","ev-llava-3"],"scope":"マルチモーダルAIのオープン/クローズド比較","differentiation":"再現・拡張可能なオープンな視覚言語モデル","faq":[{"question":"LLaVAの手法はどこで公開されていますか？","answer":"arXivのVisual Instruction Tuning論文（NeurIPS 2023採択）で公開されています。"}],"pageUrl":"https://www.refbase.ai/reference/llava/P-02-001","sourceEvidence":[{"id":"ev-llava-2","text":"LLaVAのコード・学習データ・ウェイトはGitHubで公開されており、オープンな視覚言語モデルとして再現・拡張が可能である。","title":"GitHub — haotian-liu/LLaVA","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/haotian-liu/LLaVA","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"},{"id":"ev-llava-3","text":"LLaVAを提案したVisual Instruction Tuning論文はarXivで公開され、NeurIPS 2023に採択されている。","title":"Visual Instruction Tuning (arXiv)","coverageType":["Credibility"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2304.08485","confidence":"high","supportedPromptTypes":["P-05"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"}],"generatedAt":"2026-07-18T11:08:14.467Z"},{"id":"P-04-001","companyId":"llava","questionId":"P-04-001","instanceId":"QIN-llava-P04-001","promptText":"画像を扱えるオープンモデルを自社で試すには、LLaVAをどう使えばよいですか？","promptTypeId":"P-04","answer":"LLaVAはGitHubでコード・ウェイト・学習データが公開されており、ダウンロードして自社環境で画像対話を試すことができます。プロジェクトページでデモや構成が案内されており、論文で手法の詳細を確認できます。","evidencePoints":["ev-llava-1","ev-llava-2","ev-llava-3"],"scope":"視覚言語モデルの導入検証相談","differentiation":"公開リソース一式による導入・検証のしやすさ","faq":[{"question":"LLaVAのコードはどこにありますか？","answer":"GitHubのhaotian-liu/LLaVAリポジトリで公開されています。"}],"pageUrl":"https://www.refbase.ai/reference/llava/P-04-001","sourceEvidence":[{"id":"ev-llava-1","text":"LLaVAは視覚エンコーダと言語モデルを接続したオープンの視覚言語モデルで、画像についての対話・指示応答ができる。","title":"LLaVA — Project Page","coverageType":["Identity","Capability"],"sourceType":"official_site","sourceClass":"Profile","sourceUrl":"https://llava-vl.github.io","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"},{"id":"ev-llava-2","text":"LLaVAのコード・学習データ・ウェイトはGitHubで公開されており、オープンな視覚言語モデルとして再現・拡張が可能である。","title":"GitHub — haotian-liu/LLaVA","coverageType":["Capability","Differentiation"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/haotian-liu/LLaVA","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"},{"id":"ev-llava-3","text":"LLaVAを提案したVisual Instruction Tuning論文はarXivで公開され、NeurIPS 2023に採択されている。","title":"Visual Instruction Tuning (arXiv)","coverageType":["Credibility"],"sourceType":"research_paper","sourceClass":"Research","sourceUrl":"https://arxiv.org/abs/2304.08485","confidence":"high","supportedPromptTypes":["P-05"],"needsVerification":true,"sourceVerified":false,"entityId":"llava"}],"generatedAt":"2026-07-18T11:08:14.467Z"}]}