Skip to content
local-ai

Qwen2.5-VL 7B

Alibaba · Vision-language · 7B · 33k context · Released 28 January 2025

Commercial use permitted Text + ImageText Open weights Runs on CPU Apple Silicon

A 7-billion-parameter vision-language model that reads images and documents well above its weight, with particularly strong optical character recognition and chart and table understanding. Apache 2.0, and runnable on a 12GB card.

Strengths

  • Strong OCR and document, chart, and table understanding for its size
  • Handles video as well as still images
  • Apache 2.0, and fits a 12GB card at a good quantisation

Weaknesses

  • A 7B model trails larger VLMs on hard visual reasoning
  • Vision support in local runtimes is less mature than for text-only models
  • Shorter context than the text-only Qwen models

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~6GBFits an 8GB card with short context, including the vision encoder
Q8_0~9GBA good fit for a 12GB card
FP16~17GBFull precision, needs 24GB

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~6GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

For good quality

Q8_0 · ~9GB needed

One 16GB GPU

Intel Arc A770 16GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 16GB runs →

Best quality

FP16 · ~17GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

Licence

Apache 2.0 read the licence

Benchmarks

BenchmarkScoreSourceAs of
MMMU (val)58.6 Qwen2.5-VL technical report February 2025
DocVQA (test)95.7 Qwen2.5-VL technical report (ANLS, test) February 2025

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

Qwen2.5-VL 7B: common questions

What hardware do I need to run Qwen2.5-VL 7B?
At its most compressed (Q4_K_M) it needs roughly 6GB of VRAM, and about 9GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Qwen2.5-VL 7B free for commercial use?
Yes. Qwen2.5-VL 7B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run Qwen2.5-VL 7B on Apple Silicon?
Yes. Qwen2.5-VL 7B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Qwen2.5-VL 7B run on CPU?
Yes, Qwen2.5-VL 7B can run on the CPU, though generation is slower than on a GPU.
What is Qwen2.5-VL 7B's context window?
Qwen2.5-VL 7B has a context window of 32,768 tokens, about 33k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Bartowski GGUF Q2-Q8 (imatrix)

    Vision GGUF support varies by runtime, so check the image encoder is handled.

  • Very broad coverage of models, in both static and imatrix GGUF.

  • MLX community MLX 4-bit and 8-bit

    MLX quants for Apple Silicon, usually 4-bit and 8-bit.

Recommended for

  • Local OCR, document, and screenshot understanding
  • Extracting data from charts and tables
  • A capable vision model on a 12GB card

Related models

Run it with

Glossary

Our coverage

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.