Skip to content
local-ai

Qwen3-Embedding 0.6B

Alibaba · Embedding · 0.6B · 33k context · Released 5 June 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

The smallest model in Alibaba's Qwen3-Embedding family, strong on multilingual retrieval for its size while staying light enough to run on a CPU. Apache 2.0, with user-selectable output dimensions.

Strengths

  • Strong multilingual retrieval quality for a sub-1B model
  • Selectable output dimensions, from 32 up to 1024
  • Apache 2.0, and small enough for CPU use

Weaknesses

  • The larger 4B and 8B siblings retrieve better if you have the memory
  • Newer than nomic-embed and BGE-M3, so less battle-tested in tooling
  • As always, a domain-tuned embedder may beat it on narrow content

Hardware requirements

QuantisationApprox. VRAMNotes
Q8_0~0.7GBLight enough for CPU, minimal quality loss
FP16~1.2GBFull precision

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q8_0 · ~0.7GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

Best quality

FP16 · ~1.2GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

Licence

Apache 2.0 read the licence

Benchmarks

BenchmarkScoreSourceAs of
MTEB Multilingual Mean (Task)64.33 Qwen3-Embedding-0.6B model card June 2025

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

Qwen3-Embedding 0.6B: common questions

What hardware do I need to run Qwen3-Embedding 0.6B?
At its most compressed (Q8_0) it needs roughly 0.7GB of VRAM, and about 1.2GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Qwen3-Embedding 0.6B free for commercial use?
Yes. Qwen3-Embedding 0.6B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run Qwen3-Embedding 0.6B on Apple Silicon?
Yes. Qwen3-Embedding 0.6B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Qwen3-Embedding 0.6B run on CPU?
Yes, Qwen3-Embedding 0.6B can run on the CPU, though generation is slower than on a GPU.
What is Qwen3-Embedding 0.6B's context window?
Qwen3-Embedding 0.6B has a context window of 32,768 tokens, about 33k.

Availability

Recommended for

  • Multilingual RAG on modest hardware
  • Setups that benefit from adjustable embedding dimensions
  • A current, permissively licensed embedder

Related models

Run it with

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.