ExLlamaV2
A fast inference library for running GPTQ and EXL2 quantised models on consumer NVIDIA GPUs, known for high throughput at low bit-widths.
Strengths
- Strong throughput on consumer NVIDIA cards
- The EXL2 format allows flexible, fine-grained bit-widths from 2 to 8 bits
- A dynamic generator with smart prompt caching and batching
Weaknesses
- The repo is archived: development has moved to ExLlamaV3
- NVIDIA and CUDA only
- Requires compiling C++ extensions
At a glance
- Licence
- MIT
- Pricing
- Free
- Platforms
- Windows, Linux
- GPU required
- Yes
Links
Related guides
Glossary
Entry last verified 18 August 2026.