Discover / Token & Cost Optimization
TensorRT-LLM
by NVIDIAPython
Toolkit for optimizing and deploying large language models with high performance inference.
Maturity: stable because 3y old, v1.2.1 released 105d ago. Derived from release and commit history, not a rating.
- Stars
- 14k
- Forks
- 2.6k
- Downloads / mo
- 12k
- Last commit
- 2026-08-03
- License
- NOASSERTION
- Open issues
- 1.6k
Market and trust evidence
Edition not yet matchedNo exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.
Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.
In practice
Written by AI from this repository’s README · low confidenceGetting production throughput from a large model requires kernel level optimization work most teams cannot do.
Use it when
Use it when you are serving large language or visual generation models and need specialized kernels and an efficient runtime.
Not the right pick when
The README is mostly blog links, so setup, hardware requirements and usage all live in the external documentation.
Capabilities
- specialized kernels for common inference operations
- efficient runtime for inference execution
- pythonic framework for customizing and extending the system
- documented performance and architecture guides
- published quick start examples
Cost: Free and open source
Install
Derived from the published package name in the repository, not from a model.
Video walkthroughs
🔍 AI Serving Frameworks Explained: vLLM vs TensorRT-LLM vs Ray Serve | Which One Should You Use?
TensorRT LLM 1.0 Livestream: New Easy-To-Use Pythonic Runtime
Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.
What the repository ships
Detected from the actual files in the repository root.
Latest release v1.2.1
Published 2026-04-20
Highlights
- Fixed Issue
- Fixed an issue that caused KV cache corruption (#12770)
- Infrastructure Changes
- Upgraded xgrammar and flashinfer (#12811)
Tags
README
<div align="center">
TensorRT LLM
===========================
<h4>TensorRT LLM optimizes inference for LLMs and Visual Gen models with specialized kernels for common operations, an efficient runtime, and a pythonic framework that enables you to customize and extend the system.</h4>
Architecture | Performance | Examples | Documentation | Roadmap
<div align="left">
Tech Blogs
<!-- Use github markdown link to link for the latest blog since the doc build has not happened yet. When the doc build is updated, it should be updated to the webpage link. -->
- [07/17] DeepSeek-V4 on NVIDIA Blackwell: Model-Specific and Agentic-Workload Optimizations in TensorRT LLM
✨ ➡️ link
- [07/01] Scaling Video Generation Across NVL72 Rack with TensorRT-LLM
✨ ➡️ link
- [05/15] Joint Optimization of Agent Applications and TensorRT-LLM
✨ ➡️ link
- [04/03] Tuning CUDA Graph Batch Sizes for Higher Output Throughput
✨ ➡️ link
- [04/03] DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
✨ ➡️ link
- [03/16] Optimizing MoE Communication with One-Sided AlltoAll Over NVLink
✨ ➡️ link
- [03/04] Sparse Attention in TensorRT LLM
✨ ➡️ link
- [02/06] Accelerating Long-Context Inference with Skip Softmax Attention
✨ ➡️ link
- [01/09] Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs
✨ ➡️ link
<details close>
<summary>Previous Blogs</summary>
- [10/13] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary)
✨ ➡️ link
- [09/26] Inference Time Compute Implementation in TensorRT LLM
✨ ➡️ link
- [09/19] Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly
✨ ➡️ link
- [08/29] ADP Balance Strategy
✨ ➡️ link
- [08/05] Running a High-Performance GPT-OSS-120B Inference Server with TensorRT LLM
✨ ➡️ link
- [08/01] Scaling Expert Parallelism in TensorRT LLM (Part 2: Performance Status and Optimization)
✨ ➡️ link
- [07/26] N-Gram Speculative Decoding in TensorRT LLM
✨ ➡️ link
- [06/19] Disaggregated Serving in TensorRT LLM
✨ ➡️ link
- [06/05] Scaling Expert Parallelism in TensorRT LLM (Part 1: Design and Implementation of Large-scale EP)
✨ ➡️ link
- [05/30] Optimizing DeepSeek R1 Throughput on NVIDIA Blackwell GPUs: A Deep Dive for Developers
✨ ➡️ link
- [05/23] DeepSeek R1 MTP Implementation and Optimization
✨ ➡️ link
- [05/16] Pushing Latency Boundaries: Optimizing DeepSeek-R1 Performance on NVIDIA B200 GPUs
✨ ➡️ link
</details>
Latest News
- [04/03] 🎨 TensorRT LLM now supports diffusion models for visual generation ➡️ link
<details close>
<summary>Previous News</summary>
- [08/05] 🌟 TensorRT LLM delivers Day-0 support for OpenAI's latest open-weights models: GPT-OSS-120B ➡️ link and GPT-OSS-20B ➡️ link
- [07/15] 🌟 TensorRT LLM delivers Day-0 support for LG AI Research's latest model, EXAONE 4.0 ➡️ link
- [05/22] Blackwell Breaks the 1,000 TPS/User Barrier With Meta’s Llama 4 Maverick
✨ ➡️ link
- [04/10] TensorRT LLM DeepSeek R1 performance benchmarking best practices now published.
✨ ➡️ link
- [04/05] TensorRT LLM can run Llama 4 at over 40,000 tokens per second on B200 GPUs!
- [03/22] TensorRT LLM is now fully open-source, with developments moved to GitHub!
- [03/18] 🚀🚀 NVIDIA Blackwell Delivers World-Record DeepSeek-R1 Inference Performance with TensorRT LLM ➡️ Link
- [02/28] 🌟 NAVER Place Optimizes SLM-Based Vertical Services with TensorRT LLM ➡️ Link
- [02/25] 🌟 DeepSeek-R1 performance now optimized for Blackwell ➡️ Link
- [02/20] Explore the complete guide to achieve great accuracy, high throughput, and low latency at the lowest cost for your business here.
- [02/18] Unlock #LLM inference with auto-scaling on @AWS EKS ✨ ➡️ link
- [02/12] 🦸⚡ Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling
- [02/12] 🌟 How Scaling Laws Drive Smarter, More Powerful AI
- [2025/01/25] Nvidia moves AI focus to inference cost, efficiency ➡️ link
- [2025/01/24] 🏎️ Optimize AI Inference Performance with NVIDIA Full-Stack Solutions ➡️ link
- [2025/01/23] 🚀 Fast, Low-Cost Inference Offers Key to Profitable AI ➡️ link
- [2025/01/16] Introducing New KV Cache Reuse Optimizations in TensorRT LLM ➡️ link
- [2025/01/14] 📣 Bing's Transition to LLM/SLM Models: Optimizing Search with TensorRT LLM ➡️ link
- [2025/01/04] ⚡Boost Llama 3.3 70B Inference Throughput 3x with TensorRT LLM Speculative Decoding
- [2024/12/10] ⚡ Llama 3.3 70B from AI at Meta is accelerated by TensorRT-LLM. 🌟 State-of-the-art model on par with Llama 3.1 405B for reasoning, math, instruction following and tool use. Explore the preview
- [2024/12/03] 🌟 Boost your AI inference throughput by up to 3.6x. We now support speculative decoding and tripling token throughput with our NVIDIA TensorRT-LLM. Perfect for your generative AI apps. ⚡Learn how in this technical deep dive
- [2024/12/02] Working on deploying ONNX models for performance-critical applications? Try our NVIDIA Nsight Deep Learning Designer ⚡ A user-friendly GUI and tight integration with NVIDIA TensorRT that offers:
✅ Intuitive visualization of ONNX model graphs
✅ Quick tweaking of model architecture and parameters
✅ Detailed performance profiling with either ORT or TensorRT
✅ Easy building of TensorRT engines
- [2024/11/26] 📣 Introducing TensorRT LLM for Jetson AGX Orin, making it even easier to deploy on Jetson AGX Orin with initial support in JetPack 6.1 via the v0.12.0-jetson branch of the TensorRT LLM repo. ✅ Pre-compiled TensorRT LLM wheels & containers for easy integration ✅ Comprehensive guides & docs to get you started
- [2024/11/21] NVIDIA TensorRT LLM Multiblock Attention Boosts Throughput by More Than 3x for Long Sequence Lengths on NVIDIA HGX H200
- [2024/11/19] Llama 3.2 Full-Stack Optimizations Unlock High Performance on NVIDIA GPUs
- [2024/11/09] 🚀🚀🚀 3x Faster AllReduce with NVSwitch and TensorRT LLM MultiShot
- [2024/11/09] ✨ NVIDIA advances the AI ecosystem with the AI model of LG AI Research 🙌
- [2024/11/02] 🌟🌟🌟 NVIDIA and LlamaIndex Developer Contest
🙌 Enter for a chance to win prizes including an NVIDIA® GeForce RTX™ 4080 SUPER GPU, DLI credits, and more🙌
- [2024/10/28] 🏎️🏎️🏎️ NVIDIA GH200 Superchip Accelerates Inference by 2x in Multiturn Interactions with Llama Models
- [2024/10/22] New 📝 Step-by-step instructions on how to
✅ Optimize LLMs with NVIDIA TensorRT-LLM,
✅ Deploy the optimized models with Triton Inference Server,
✅ Autoscale LLMs deployment in a Kubernetes environment.
🙌 Technical Deep Dive:
- [2024/10/07] 🚀🚀🚀Optimizing Microsoft Bing Visual Search with NVIDIA Accelerated Libraries
- [2024/09
Truncated. Read the full README on GitHub ↗