Skip to content
local-ai

Text Generation Inference (TGI)

Orchestration Advanced Open source Free

Hugging Face's production-grade Rust and Python server for serving large language models with optimised, high-throughput inference.

Strengths

  • Production features: continuous batching, tensor parallelism, token streaming, metrics
  • Broad quantisation support (bitsandbytes, GPTQ, AWQ, Marlin, FP8, EETQ)
  • OpenAI Chat Completions-compatible Messages API

Weaknesses

  • In maintenance mode as of March 2026, with vLLM and SGLang recommended as successors
  • GPU dependency makes CPU deployment impractical
  • Heavy install and multi-GPU setup complexity

At a glance

Licence
Apache-2.0
Pricing
Free
Platforms
Linux, Docker
GPU required
Yes

Links

Related guides

See also

Glossary

Entry last verified 18 August 2026.