Ollama
Ollama handles downloading, storing, and serving models behind a single command and an OpenAI-compatible endpoint, which is why so many desktop apps and coding tools integrate with it. It trades some of llama.cpp's fine control for convenience, which is the right trade for most people getting started.
Strengths
- Genuinely simple to install and run a first model
- A local API that a large ecosystem of tools already targets
- Sensible defaults and painless model management
Weaknesses
- Less direct control over inference parameters than raw llama.cpp
- Defaults can hide what is happening, which slows learning the internals
- Getting non-default quantisations or custom models in takes extra steps
At a glance
- Licence
- MIT
- Pricing
- Free
- Platforms
- macOS, Windows, Linux, Docker
- GPU required
- No
Links
Related guides
See also
Glossary
Our coverage
- Apple's M5 Mac Studio and Mac mini refresh puts 512GB of unified memory on the desk
- Qwen3.8 arrives in the open: a runnable 27B, and a text-only 2.4T flagship
- Meta releases Muse Glimmer, a 30B open model under Apache 2.0 built to run on one GPU
- Ollama's August update speeds up Apple Silicon inference with speculative decoding
- Alibaba releases Qwen3.8-Max as open weights, with a runnable 27B to follow
- Mistral releases the Mistral 3 family under Apache 2.0
- Qwen3.6 arrives with new mixture-of-experts checkpoints
- Strix Halo mini PCs put 128GB of unified memory within reach for local AI
- Alibaba releases Qwen3.5, extending its open-weight line
- NVIDIA's DGX Spark launches, a 128GB desktop AI machine
- Meta releases Llama 3.3 70B, matching much larger models at lower cost
Entry last verified 30 July 2026.