llama.cpp
llama.cpp introduced the GGUF format and remains the reference for running quantised models on consumer hardware. Many higher-level tools, Ollama among them, wrap it. Working with it directly gives you control over context length, KV cache handling, offloading, and sampling that convenience wrappers abstract away.
Strengths
- Runs well across a wide range of hardware, including CPU-only
- Fine-grained control over quantisation, context, and offloading
- The GGUF ecosystem and community support are enormous
Weaknesses
- The learning curve is real, especially the build and flag options
- You manage models and updates yourself, with less hand-holding
- Documentation assumes a fair amount of command-line familiarity
At a glance
- Licence
- MIT
- Pricing
- Free
- Platforms
- macOS, Windows, Linux
- GPU required
- No
Links
Works with
Related guides
See also
Glossary
Our coverage
- Several open video models exclude UK users by licence, outputs included
- Apple's M5 Mac Studio and Mac mini refresh puts 512GB of unified memory on the desk
- Liquid AI releases LFM2.5-VL-3B, an open-weight vision model built to run on-device
- Meta releases Muse Glimmer, a 30B open model under Apache 2.0 built to run on one GPU
- Ollama's August update speeds up Apple Silicon inference with speculative decoding
- Alibaba releases Qwen3.8-Max as open weights, with a runnable 27B to follow
- Mistral releases the Mistral 3 family under Apache 2.0
- DeepSeek releases V4 as open weights under MIT
- Qwen3.6 arrives with new mixture-of-experts checkpoints
- Strix Halo mini PCs put 128GB of unified memory within reach for local AI
- Alibaba releases Qwen3.5, extending its open-weight line
- Meta releases Llama 3.3 70B, matching much larger models at lower cost
Entry last verified 30 July 2026.