llama-swap
A lightweight Go proxy that hot-swaps models behind an OpenAI-compatible endpoint, commonly used in front of llama.cpp to serve many models from one machine.
Strengths
- Single Go binary, zero external dependencies, minimal overhead
- Works with llama.cpp, vLLM, whisper.cpp, and ComfyUI backends
- Automatic model unloading via a configurable idle timeout
Weaknesses
- Not an inference engine itself, so it needs a backend server to be useful
- Config-driven YAML setup for each model
- Small single-maintainer project
At a glance
- Licence
- MIT
- Pricing
- Free
- Platforms
- macOS, Windows, Linux, Docker
- GPU required
- No
Links
Related guides
Glossary
Entry last verified 18 August 2026.