vLLM
vLLM is aimed at serving rather than casual local use. It keeps GPUs busy with continuous batching and efficient memory management, exposes an OpenAI-compatible API, and supports a wide range of models and hardware. If Ollama is how you try a model, vLLM is often how you serve it to a team.
Strengths
- High throughput and concurrency, ideal for multi-user serving
- Broad model and hardware coverage
- OpenAI-compatible API, so a drop-in for many applications
Weaknesses
- Heavier to set up and operate than a desktop app like Ollama
- Overkill for a single user on one machine
- Gets the most from data-centre GPUs, and needs a GPU to be worthwhile
At a glance
- Licence
- Apache 2.0
- Pricing
- Free
- Platforms
- Linux, Docker
- GPU required
- Yes
Links
Works with
Related guides
Glossary
Our coverage
- DeepSeek releases V4.1 Flash, an MIT-licensed multimodal model with an unusual architecture
- NVIDIA upstreams llama.cpp optimisations it says make local inference up to 1.9x faster
- GLM-5.3 open weights land, under a bespoke licence rather than MIT
- Qwen3.8 arrives in the open: a runnable 27B, and a text-only 2.4T flagship
- Liquid AI releases LFM2.5-VL-3B, an open-weight vision model built to run on-device
- NVIDIA releases Nemotron 3.5 Lightning, a fully open 30B mixture-of-experts
- Moonshot releases Kimi K3, a 2.8-trillion-parameter open model for server-class self-hosting
Entry last verified 30 July 2026.