Skip to content
local-ai

ExLlamaV2

Inference engines Advanced Open source Free

A fast inference library for running GPTQ and EXL2 quantised models on consumer NVIDIA GPUs, known for high throughput at low bit-widths.

Strengths

  • Strong throughput on consumer NVIDIA cards
  • The EXL2 format allows flexible, fine-grained bit-widths from 2 to 8 bits
  • A dynamic generator with smart prompt caching and batching

Weaknesses

  • The repo is archived: development has moved to ExLlamaV3
  • NVIDIA and CUDA only
  • Requires compiling C++ extensions

At a glance

Licence
MIT
Pricing
Free
Platforms
Windows, Linux
GPU required
Yes

Links

Related guides

See also

Glossary

Entry last verified 18 August 2026.