TensorRT-LLM
NVIDIA's library for extracting peak inference performance from NVIDIA GPUs. It compiles optimised engines for a model and hardware pairing, delivering the best throughput and latency available on supported cards, at the cost of setup complexity.
Strengths
- Peak throughput and latency on NVIDIA GPUs
- Production-grade optimisations, including quantisation and in-flight batching
- Integrates with NVIDIA Triton for multi-model serving
Weaknesses
- NVIDIA hardware only
- Complex to set up, with a model compilation step for each configuration
- Overkill outside production or performance-critical deployments
At a glance
- Licence
- Apache 2.0
- Pricing
- Free
- Platforms
- Linux, Docker
- GPU required
- Yes
Links
Entry last verified 30 July 2026.