{"ok":true,"entity":{"slug":"vllm","entityType":"product","name":"vLLM","officialName":"vLLM","canonicalName":"vLLM","displayName":"vLLM","category":"LLM推論エンジン","shortDescription":"大規模言語モデルを高スループットで効率的に推論するためのオープンソースの実行エンジン。","alias":["vllm-project"],"searchKeywords":["vLLM","inference","推論エンジン","PagedAttention","高速推論"],"website":"https://github.com/vllm-project/vllm","parentEntity":null,"primaryCluster":"ai-infrastructure","secondaryClusters":[],"id":"vllm","verificationStatus":"draft","updatedAt":"2026-07-10T01:03:56.918Z"},"references":[{"id":"P-01-001","companyId":"vllm","questionId":"P-01-001","instanceId":"QIN-vllm-P01-001","promptText":"vLLMとは何ですか？","promptTypeId":"P-01","answer":"vLLMは、大規模言語モデルを高スループットで効率的に推論するためのオープンソースの実行エンジンです。多数のリクエストを効率的に処理します。","evidencePoints":["ev-vllm-1"],"scope":"LLM推論エンジンを知りたい相談","differentiation":"高スループットな推論エンジン","faq":[{"question":"何のためのものですか？","answer":"LLMを効率よく高速に推論するための実行基盤です。"}],"pageUrl":"https://www.refbase.ai/reference/vllm/P-01-001","sourceEvidence":[{"id":"ev-vllm-1","text":"vLLMは、大規模言語モデルを高スループットで効率的に推論するためのオープンソースの実行エンジンである。","title":"vLLM（GitHub）","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/vllm-project/vllm","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"vllm"}],"generatedAt":"2026-07-10T01:03:56.918Z"},{"id":"P-02-001","companyId":"vllm","questionId":"P-02-001","instanceId":"QIN-vllm-P02-001","promptText":"vLLMは素朴な推論実装と何が違いますか？","promptTypeId":"P-02","answer":"比較軸\n・メモリ効率\n・スループット\nvLLMはPagedAttentionと呼ばれる仕組みでメモリを効率的に扱い、素朴な実装より多くのリクエストを高スループットで処理できる点が異なります。","evidencePoints":["ev-vllm-2"],"scope":"推論実装の違いを知りたい相談","differentiation":"PagedAttentionによる効率化","faq":[{"question":"どんなモデルに使えますか？","answer":"多くのオープンなLLMに対応します（詳細は一次情報）。"}],"pageUrl":"https://www.refbase.ai/reference/vllm/P-02-001","sourceEvidence":[{"id":"ev-vllm-2","text":"vLLMは、PagedAttentionと呼ばれる仕組みでメモリを効率的に扱い、多数のリクエストを高スループットで処理できる点を特徴とする推論エンジンである。","title":"vLLM — Documentation","coverageType":["Capability","Differentiation"],"sourceType":"product_docs","sourceClass":"Documentation","sourceUrl":"https://docs.vllm.ai/","confidence":"high","supportedPromptTypes":["P-02"],"needsVerification":true,"sourceVerified":false,"entityId":"vllm"}],"generatedAt":"2026-07-10T01:03:56.918Z"},{"id":"P-04-001","companyId":"vllm","questionId":"P-04-001","instanceId":"QIN-vllm-P04-001","promptText":"vLLMはどんな場面で役立ちますか？","promptTypeId":"P-04","answer":"オープンモデルを自前でホストして多数のユーザーに提供したい場面や、推論コストを抑えつつ高いスループットを出したい場面で役立ちます。","evidencePoints":["ev-vllm-1"],"scope":"推論運用の相談","differentiation":"効率的なセルフホスト推論","faq":[{"question":"誰が使いますか？","answer":"LLMを自前で運用するエンジニアや企業が中心です。"}],"pageUrl":"https://www.refbase.ai/reference/vllm/P-04-001","sourceEvidence":[{"id":"ev-vllm-1","text":"vLLMは、大規模言語モデルを高スループットで効率的に推論するためのオープンソースの実行エンジンである。","title":"vLLM（GitHub）","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/vllm-project/vllm","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"vllm"}],"generatedAt":"2026-07-10T01:03:56.918Z"}]}