Skip to content
local-ai

Qwen2.5-VL 7B

Alibaba · Vision-language · 7B · 33k context · Released 28 January 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

A 7-billion-parameter vision-language model that reads images and documents well above its weight, with particularly strong optical character recognition and chart and table understanding. Apache 2.0, and runnable on a 12GB card.

Strengths

  • Strong OCR and document, chart, and table understanding for its size
  • Handles video as well as still images
  • Apache 2.0, and fits a 12GB card at a good quantisation

Weaknesses

  • A 7B model trails larger VLMs on hard visual reasoning
  • Vision support in local runtimes is less mature than for text-only models
  • Shorter context than the text-only Qwen models

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~6GBFits an 8GB card with short context, including the vision encoder
Q8_0~9GBA good fit for a 12GB card
FP16~17GBFull precision, needs 24GB

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Apache 2.0 read the licence

Availability

Recommended for

  • Local OCR, document, and screenshot understanding
  • Extracting data from charts and tables
  • A capable vision model on a 12GB card

Run it with

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.