Llama 3.2 11B Vision
Meta · Vision-language · 11B · 131k context · Released 25 September 2024
Permitted with conditions
Open weights
Runs on CPU
Apple Silicon
Meta's 11-billion-parameter vision-language model, a solid general choice for understanding images alongside text on a mid-range card. Well supported across local tooling, with a long 128k context window.
Strengths
- Capable general image-and-text understanding on 12GB to 16GB cards
- Broad local tooling support, including Ollama
- Long 128k context window
Weaknesses
- Weaker at dense OCR and documents than Qwen2.5-VL of similar size
- The Llama licence carries conditions, unlike Apache 2.0 models
- Needs 12GB or more for usable quality
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~8GB | Fits a 12GB card with short context, including the vision encoder |
| Q8_0 | ~13GB | A good fit for a 16GB card |
| FP16 | ~22GB | Full precision, needs 24GB |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Llama 3.2 Community License — read the licence
Availability
- Official page
- Hugging Face
- ollama run llama3.2-vision:11b
Recommended for
- General image understanding on a 12GB to 16GB card
- Users already in the Llama ecosystem who want vision
- A well-supported starting point for local multimodal work
Run it with
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.