Qwen3-Embedding 0.6B
Alibaba · Embedding · 0.6B · 33k context · Released 5 June 2025
The smallest model in Alibaba's Qwen3-Embedding family, strong on multilingual retrieval for its size while staying light enough to run on a CPU. Apache 2.0, with user-selectable output dimensions.
Strengths
- Strong multilingual retrieval quality for a sub-1B model
- Selectable output dimensions, from 32 up to 1024
- Apache 2.0, and small enough for CPU use
Weaknesses
- The larger 4B and 8B siblings retrieve better if you have the memory
- Newer than nomic-embed and BGE-M3, so less battle-tested in tooling
- As always, a domain-tuned embedder may beat it on narrow content
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q8_0 | ~0.7GB | Light enough for CPU, minimal quality loss |
| FP16 | ~1.2GB | Full precision |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
Q8_0 · ~0.7GB needed
One 12GB GPU
NVIDIA GeForce RTX 3060 12GBor a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
At least 32GB of system RAM alongside the card.
around £700–£1,100
What else 12GB runs →Best quality
FP16 · ~1.2GB needed
One 12GB GPU
NVIDIA GeForce RTX 3060 12GBor a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
At least 32GB of system RAM alongside the card.
around £700–£1,100
What else 12GB runs →Licence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| MTEB Multilingual Mean (Task) | 64.33 | Qwen3-Embedding-0.6B model card | June 2025 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
MTEB Multilingual Mean (Task)
higher is better- Qwen3-Embedding 0.6B 64.33
Qwen3-Embedding-0.6B model card · June 2025
- BGE-M3 59.56
Qwen3-Embedding-0.6B model card (comparison table, reported by Qwen) · June 2025
Qwen3-Embedding 0.6B: common questions
- What hardware do I need to run Qwen3-Embedding 0.6B?
- At its most compressed (Q8_0) it needs roughly 0.7GB of VRAM, and about 1.2GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is Qwen3-Embedding 0.6B free for commercial use?
- Yes. Qwen3-Embedding 0.6B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
- Can I run Qwen3-Embedding 0.6B on Apple Silicon?
- Yes. Qwen3-Embedding 0.6B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- Does Qwen3-Embedding 0.6B run on CPU?
- Yes, Qwen3-Embedding 0.6B can run on the CPU, though generation is slower than on a GPU.
- What is Qwen3-Embedding 0.6B's context window?
- Qwen3-Embedding 0.6B has a context window of 32,768 tokens, about 33k.
Availability
Recommended for
- Multilingual RAG on modest hardware
- Setups that benefit from adjustable embedding dimensions
- A current, permissively licensed embedder
Related models
Run it with
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.