Skip to content
local-ai

Llama 3.2 11B Vision

Meta · Vision-language · 11B · 131k context · Released 25 September 2024

Permitted with conditions Open weights Runs on CPU Apple Silicon

Meta's 11-billion-parameter vision-language model, a solid general choice for understanding images alongside text on a mid-range card. Well supported across local tooling, with a long 128k context window.

Strengths

  • Capable general image-and-text understanding on 12GB to 16GB cards
  • Broad local tooling support, including Ollama
  • Long 128k context window

Weaknesses

  • Weaker at dense OCR and documents than Qwen2.5-VL of similar size
  • The Llama licence carries conditions, unlike Apache 2.0 models
  • Needs 12GB or more for usable quality

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~8GBFits a 12GB card with short context, including the vision encoder
Q8_0~13GBA good fit for a 16GB card
FP16~22GBFull precision, needs 24GB

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Llama 3.2 Community License read the licence

Availability

Recommended for

  • General image understanding on a 12GB to 16GB card
  • Users already in the Llama ecosystem who want vision
  • A well-supported starting point for local multimodal work

Run it with

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.