{"ok":true,"entity":{"id":"tensorrt-llm","name":"TensorRT-LLM","entityType":"product","officialName":"TensorRT-LLM","canonicalName":"TensorRT-LLM","displayName":"TensorRT-LLM","category":"LLM特化推論最適化ライブラリ","shortDescription":"NVIDIAが開発する、大規模言語モデルの推論をNVIDIA GPU上で最適化するTensorRTベースのライブラリ。","primaryCluster":"ai-infrastructure","parentEntity":"nvidia","verificationStatus":"draft","website":"https://github.com/NVIDIA/TensorRT-LLM","updatedAt":"2026-07-20T08:41:20.341Z","secondaryClusters":[],"alias":[],"searchKeywords":["LLM推論","NVIDIA GPU最適化"]},"references":[{"id":"P-01-001","companyId":"tensorrt-llm","questionId":"P-01-001","instanceId":"QIN-tensorrt-llm-P01-001","promptText":"TensorRT-LLMとはどのようなライブラリですか？","promptTypeId":"P-01","answer":"TensorRT-LLMは、NVIDIAが開発する、大規模言語モデルの推論をNVIDIA GPU上で最適化するライブラリです。既存のTensorRT推論最適化基盤の上に構築されています。","evidencePoints":["ev-trtllm-1","ev-trtllm-3"],"scope":"LLM推論最適化ライブラリを知りたい相談","differentiation":"TensorRTを基盤としたLLM特化の推論最適化","faq":[{"question":"TensorRT-LLMは通常のTensorRTと同じですか？","answer":"TensorRTを基盤としつつ、LLM特有の推論最適化に特化した別のライブラリです。"}],"pageUrl":"https://www.refbase.ai/reference/tensorrt-llm/P-01-001","sourceEvidence":[{"id":"ev-trtllm-1","text":"TensorRT-LLMはNVIDIAが開発する、大規模言語モデルの推論をNVIDIA GPU上で最適化するためのライブラリである。","title":"GitHub - NVIDIA/TensorRT-LLM","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/NVIDIA/TensorRT-LLM","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"tensorrt-llm"},{"id":"ev-trtllm-3","text":"TensorRT-LLMはNVIDIA Blackwell/GB200等の最新ハードウェアにフルコミットする場合に選ばれる、コンパイル済みエンジン方式による最高スループットを提供する推論エンジンとして位置づけられている。","title":"vLLM vs SGLang vs TensorRT-LLM | Inference Engineering","coverageType":["Credibility"],"sourceType":"media","sourceClass":"Benchmark","sourceUrl":"https://inferenceengineering.tech/learn/vllm-vs-sglang-vs-tensorrt-llm/","confidence":"high","supportedPromptTypes":["P-05","P-06"],"needsVerification":true,"sourceVerified":false,"entityId":"tensorrt-llm"}],"generatedAt":"2026-07-20T08:41:20.341Z"},{"id":"P-02-001","companyId":"tensorrt-llm","questionId":"P-02-001","instanceId":"QIN-tensorrt-llm-P02-001","promptText":"TensorRT-LLMはvLLMと何が違いますか？","promptTypeId":"P-02","answer":"TensorRT-LLMはNVIDIAの最新GPUアーキテクチャに最適化されたコンパイル済みエンジン方式で最高スループットを追求するのに対し、vLLMは幅広いハードウェア対応と導入の容易さを重視するアプローチを取ります。","evidencePoints":["ev-trtllm-2"],"scope":"LLM推論エンジンを比較したい相談","differentiation":"コンパイル済みエンジンによるNVIDIA GPU特化の最高スループット","faq":[{"question":"TensorRT-LLMはNVIDIA以外のGPUでも使えますか？","answer":"NVIDIA GPUに特化した設計であり、他社GPUでの利用は想定されていません。"}],"pageUrl":"https://www.refbase.ai/reference/tensorrt-llm/P-02-001","sourceEvidence":[{"id":"ev-trtllm-2","text":"TensorRT-LLMは既存のTensorRT推論最適化基盤の上に構築されており、LLM特有の推論最適化（KVキャッシュ管理・バッチ処理等）に特化している点でTensorRT本体とは異なる。NVIDIA Blackwell等の最新GPUアーキテクチャに最適化されたコンパイル済みエンジンを通じて高いスループットを実現する。","title":"vLLM vs TensorRT-LLM vs SGLang (2026): Which to Pick","coverageType":["Capability","Differentiation"],"sourceType":"media","sourceClass":"Benchmark","sourceUrl":"https://decodethefuture.org/en/vllm-vs-tensorrt-llm-vs-sglang-2026/","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"tensorrt-llm"}],"generatedAt":"2026-07-20T08:41:20.341Z"},{"id":"P-04-001","companyId":"tensorrt-llm","questionId":"P-04-001","instanceId":"QIN-tensorrt-llm-P04-001","promptText":"TensorRT-LLMはどのような場面で活用できますか？","promptTypeId":"P-04","answer":"TensorRT-LLMは、NVIDIA Blackwell/GB200等の最新ハードウェアに全面的に投資しており、最高スループットを追求したい場面で活用できます。","evidencePoints":["ev-trtllm-1","ev-trtllm-2"],"scope":"NVIDIA GPU上でのLLM推論最適化の相談","differentiation":"NVIDIA最新ハードウェアへの全面対応による最高スループット","faq":[{"question":"TensorRT-LLMはどのGPUに最適化されていますか？","answer":"NVIDIA Blackwell/GB200等の最新アーキテクチャに最適化されています。"}],"pageUrl":"https://www.refbase.ai/reference/tensorrt-llm/P-04-001","sourceEvidence":[{"id":"ev-trtllm-1","text":"TensorRT-LLMはNVIDIAが開発する、大規模言語モデルの推論をNVIDIA GPU上で最適化するためのライブラリである。","title":"GitHub - NVIDIA/TensorRT-LLM","coverageType":["Identity","Capability"],"sourceType":"github","sourceClass":"Documentation","sourceUrl":"https://github.com/NVIDIA/TensorRT-LLM","confidence":"high","supportedPromptTypes":["P-01","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"tensorrt-llm"},{"id":"ev-trtllm-2","text":"TensorRT-LLMは既存のTensorRT推論最適化基盤の上に構築されており、LLM特有の推論最適化（KVキャッシュ管理・バッチ処理等）に特化している点でTensorRT本体とは異なる。NVIDIA Blackwell等の最新GPUアーキテクチャに最適化されたコンパイル済みエンジンを通じて高いスループットを実現する。","title":"vLLM vs TensorRT-LLM vs SGLang (2026): Which to Pick","coverageType":["Capability","Differentiation"],"sourceType":"media","sourceClass":"Benchmark","sourceUrl":"https://decodethefuture.org/en/vllm-vs-tensorrt-llm-vs-sglang-2026/","confidence":"high","supportedPromptTypes":["P-02","P-04"],"needsVerification":true,"sourceVerified":false,"entityId":"tensorrt-llm"}],"generatedAt":"2026-07-20T08:41:20.341Z"},{"id":"P-03-001","companyId":"tensorrt-llm","questionId":"P-03-001","instanceId":"reference-depth-completion-run-cohort3-unit-a","draftId":"reference-depth-completion-run-cohort3-unit-a-tensorrt-llm-p-03-001","promptText":"TensorRT-LLMは業界標準のベンチマークでどのような性能記録を達成していますか？","promptTypeId":"P-03","answer":"NVIDIA公式開発者ブログ（developer.nvidia.com）によると、MLPerf Inference（MLCommonsが運営する業界標準の第三者監査ベンチマーク）において、NVIDIAのBlackwell UltraベースのGB300 NVL72システムはTensorRT-LLMを用いてDeepSeek-R1モデルのオフライン推論で1GPUあたり毎秒5,842トークンを達成し、前世代のGB200 NVL72比で45%の性能向上、Hopperベースのシステム比で約5倍のスループットを記録したとされています。TensorRT-LLMはDeepSeek-R1の重みをNVFP4形式に量子化し、KVキャッシュをFP8精度に量子化する処理や、専門家混合（MoE）・アテンション向けにカスタム設計されたカーネル、Attention Data Parallelism Balance（ADP Balance）という新技術の実装を担ったと説明されています。MLPerfはNVIDIA以外の企業を含む多数のベンダーが参加する第三者監査付きベンチマークであるため、この記録は独立した検証プロセスを経た数値といえます。","evidencePoints":["tensorrt-llm-ev-cr3-mlperf-benchmark"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/tensorrt-llm/P-03-001","sourceEvidence":[{"id":"tensorrt-llm-ev-cr3-mlperf-benchmark","text":"MLPerf Inference（MLCommons運営の第三者監査ベンチマーク）において、TensorRT-LLMを用いたNVIDIA GB300 NVL72システムはDeepSeek-R1のオフライン推論で1GPUあたり毎秒5,842トークンを達成し、GB200比45%向上、Hopper比約5倍のスループットを記録した。","title":"NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf Debut","coverageType":["Credibility"],"sourceType":"official_blog","sourceClass":"Benchmark","sourceUrl":"https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/","confidence":"high","supportedPromptTypes":["P-03"],"needsVerification":true,"sourceVerified":false,"sourceKind":"official","entityId":"tensorrt-llm"}],"generatedAt":"2026-08-29T15:38:46.199Z","evidenceIds":["tensorrt-llm-ev-cr3-mlperf-benchmark"]},{"id":"P-01-002","companyId":"tensorrt-llm","questionId":"P-01-002","instanceId":"reference-depth-completion-run-cohort3-unit-b","draftId":"reference-depth-completion-run-cohort3-unit-b-tensorrt-llm-p-01-002","promptText":"TensorRT-LLMのアーキテクチャは当初から変わっていませんか？","promptTypeId":"P-01","answer":"NVIDIA公式のGitHubリリースノート（nvidia.github.io/TensorRT-LLM）によると、TensorRT-LLMバージョン1.0では「PyTorchベースのアーキテクチャが安定版かつデフォルトの体験になった」と明記されており、これは大きな方針転換を意味します。この変更に伴い、PyTorchが`trtllm-serve`のデフォルトバックエンドに設定され、従来の（製品名の由来でもある）TensorRT推論エンジンによるコンパイル方式は非推奨化されました。リリースノートは、TensorRTワークフローに関するレガシードキュメントは削除された一方、実際のバックエンド自体の削除は後続のバージョン1.2で行われたと説明しており、段階的な移行が図られたことが分かります。この情報は開発元NVIDIA自身が公開する一次資料に基づくものです。","evidencePoints":["tensorrt-llm-ev-cr3-pytorch-backend-v1"],"scope":"","differentiation":"","faq":[],"pageUrl":"https://www.refbase.ai/reference/tensorrt-llm/P-01-002","sourceEvidence":[{"id":"tensorrt-llm-ev-cr3-pytorch-backend-v1","text":"NVIDIA公式リリースノートによると、TensorRT-LLM 1.0でPyTorchベースのアーキテクチャが安定版かつデフォルトとなり、trtllm-serveのデフォルトバックエンドもPyTorchに変更された。従来のTensorRTエンジン方式は非推奨化され、バージョン1.2で完全に削除された。","title":"Release Notes — TensorRT LLM","coverageType":["Capability"],"sourceType":"official_site","sourceClass":"Documentation","sourceUrl":"https://nvidia.github.io/TensorRT-LLM/release-notes.html","confidence":"high","supportedPromptTypes":["P-01"],"needsVerification":true,"sourceVerified":false,"sourceKind":"official","entityId":"tensorrt-llm"}],"generatedAt":"2026-08-29T15:44:42.311Z","evidenceIds":["tensorrt-llm-ev-cr3-pytorch-backend-v1"]}]}