Text Generation Inference (TGI)
Hugging Face's production-grade Rust and Python server for serving large language models with optimised, high-throughput inference.
Strengths
- Production features: continuous batching, tensor parallelism, token streaming, metrics
- Broad quantisation support (bitsandbytes, GPTQ, AWQ, Marlin, FP8, EETQ)
- OpenAI Chat Completions-compatible Messages API
Weaknesses
- In maintenance mode as of March 2026, with vLLM and SGLang recommended as successors
- GPU dependency makes CPU deployment impractical
- Heavy install and multi-GPU setup complexity
At a glance
- Licence
- Apache-2.0
- Pricing
- Free
- Platforms
- Linux, Docker
- GPU required
- Yes
Links
Related guides
Glossary
Entry last verified 18 August 2026.