Model rankings
A curated summary of well-regarded models by task and size, drawn from external benchmarks. We do not run our own benchmarks; every entry is attributed and dated so you can judge it yourself.
General chat
General-purpose assistants and instruction following.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Medium | Qwen3 32B Alibaba An independent aggregate index. The strongest general model that fits a single 24GB card in this set. | Artificial Analysis Intelligence Index v4.1.1: 11 | Artificial Analysis | August 2026 |
| Medium | Mistral Small 3.2 24B Mistral AI A releasing-organisation figure. A permissive, multimodal European alternative on a 24GB card. | HumanEval Plus (Pass@5) 92.9% | Mistral Small 3.2 model card | June 2025 |
| Large | Llama 3.3 70B Instruct Meta A releasing-organisation figure. Treat as a claim pending independent confirmation. | MMLU 86.0 | Meta model card | December 2024 |
| Large | gpt-oss-120b OpenAI An independent aggregate index across nine evaluations. Runs on a single 80GB card. | Artificial Analysis Intelligence Index v4.1.1: 24 | Artificial Analysis | August 2026 |
Coding
Code generation, completion, and understanding.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Qwen2.5-Coder 7B Instruct Alibaba A releasing-organisation figure. HumanEval measures single-function generation, not agentic coding. | HumanEval 88.4% | Qwen2.5-Coder technical report | September 2024 |
| Small | Devstral Small Mistral AI A releasing-organisation figure. SWE-bench measures agentic, multi-file software tasks. | SWE-bench Verified 53.6% | Mistral AI model card (Devstral Small 1.1) | July 2025 |
| Medium | Qwen3-Coder 30B-A3B Alibaba A releasing-organisation figure; reproducibility is scaffold-sensitive. A mixture-of-experts coder that runs in roughly 18GB. | SWE-bench Verified 51.6 | Qwen (official repo, OpenHands scaffold) | August 2025 |
| Medium | Qwen2.5-Coder 32B Alibaba A releasing-organisation figure. The strongest local coder that fits a 24GB card. | HumanEval 92.7, Aider Pass@2 73.7 | Qwen2.5-Coder technical report / blog | November 2024 |
| Large | gpt-oss-120b OpenAI An independent leaderboard figure. Aider Polyglot measures multi-language code editing. | Aider Polyglot 41.8% | Aider polyglot leaderboard | August 2026 |
| Very large | DeepSeek-R1 DeepSeek An independent leaderboard figure. Server-class hardware to run. | Aider Polyglot 56.9% | Aider polyglot leaderboard | August 2026 |
Reasoning
Maths, logic, and multi-step problem solving.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Qwen3 8B Alibaba A releasing-organisation figure. Strong reasoning for a model that fits 8GB. | GPQA Diamond 62.0, AIME 2024 76.0 | Qwen3 technical report (thinking mode) | May 2025 |
| Small | Qwen3 14B Alibaba A releasing-organisation figure with thinking mode enabled. | GPQA Diamond 64.0, AIME 2024 79.3 | Qwen3 technical report (thinking mode) | May 2025 |
| Medium | QwQ 32B Alibaba A releasing-organisation figure. Treat as a claim pending independent confirmation. | MMLU-Pro 69.07 | Qwen model card | March 2025 |
| Medium | DeepSeek-R1-Distill-Qwen 32B DeepSeek A releasing-organisation figure. AIME is a hard maths benchmark; strong here does not imply strong everywhere. | AIME 2024 72.6% | DeepSeek model card | January 2025 |
| Large | gpt-oss-120b OpenAI A releasing-organisation figure at high reasoning effort. | GPQA Diamond 80.1, AIME 2024 95.8 | OpenAI gpt-oss model card (high reasoning) | August 2025 |
| Very large | DeepSeek-R1 DeepSeek A releasing-organisation figure. Server-class hardware to run; most people use a distill. | MMLU 90.8, AIME 2024 79.8% | DeepSeek model card | January 2025 |
Embedding
Text embeddings for retrieval and RAG.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Qwen3 Embedding 0.6B Alibaba A releasing-organisation figure. Tiny and fast; the 8B sibling scores higher. | MTEB Multilingual Mean (Task) 64.33 | Qwen3-Embedding model card | June 2025 |
| Small | Qwen3-Embedding 8B Alibaba A leaderboard ranking rather than a self-reported score. The lighter 0.6B version of this family is in our catalogue. | MTEB Multilingual 70.58 (No.1) | Qwen, MTEB multilingual leaderboard | June 2025 |
Reranking
Re-scoring retrieved passages to sharpen RAG results.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Qwen3-Reranker 0.6B Alibaba A releasing-organisation figure. Larger 4B and 8B siblings score higher. | MTEB-R 65.80 | Qwen3-Reranker model card | June 2025 |
| Tiny | BGE-reranker-v2-m3 BAAI A mature, widely-supported default. Figure measured by a competitor, not BAAI. | MTEB-R 57.03 | Qwen3-Reranker card comparison (competitor-measured) | June 2025 |
Vision-language
Understanding images alongside text.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Qwen2.5-VL 7B Alibaba A releasing-organisation figure. Especially strong on documents and OCR. | DocVQA (test) 95.7, MMMU (val) 58.6 | Qwen2.5-VL technical report | February 2025 |
| Small | Llama 3.2 11B Vision Meta A releasing-organisation figure (MMMU reported with chain-of-thought). | MMMU (val) 50.7, DocVQA (test) 88.4 | Meta Llama 3.2 Vision model card | September 2024 |
Image generation
Generating images from text prompts.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Speech
Speech recognition and synthesis.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Tiny | Whisper large-v3 OpenAI An independent leaderboard figure; word error rate, so lower is better. | Open ASR Leaderboard average WER 7.44 (lower is better) | Open ASR Leaderboard | October 2025 |
Head-to-head on shared benchmarks
Where two or more models we cover report the same benchmark, we line the scores up. Most are developer-reported, from the source and date on each row, so testing conditions differ between them. Scores are comparable within a board, not across boards.
General chat
Aider Polyglot
higher is better- DeepSeek-R1 56.9%
Aider polyglot leaderboard (independent) · August 2026
- gpt-oss-120b 41.8%
Aider polyglot leaderboard (independent) · August 2026
- Qwen3 32B 40.0%
Aider polyglot leaderboard (independent) · August 2026
AIME 2025
higher is better- gpt-oss-120b 92.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 91.7
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 72.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 70.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 14B 70.4
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 69.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 8B 67.3
Qwen3 technical report (thinking mode) · May 2025
-
Artificial Analysis Intelligence Index v4.1.1
higher is better- gpt-oss-120b 24
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent) · August 2026
- Qwen3 32B 11
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent, figure marked estimated) · August 2026
-
Artificial Analysis (independent) · August 2026
DocVQA (test)
higher is better- Qwen2.5-VL 7B 95.7
Qwen2.5-VL technical report (ANLS, test) · February 2025
- Llama 3.2 11B Vision 88.4
Meta Llama 3.2 Vision model card (ANLS, test) · September 2024
LiveCodeBench v5
higher is better- Qwen3 32B 65.7
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 62.6
Qwen3 technical report (thinking mode) · May 2025
MBPP
higher is better-
Qwen2.5-Coder technical report (Table 16) · September 2024
-
Qwen2.5-Coder technical report · November 2024
MMLU
higher is better- DeepSeek-R1 90.8
DeepSeek model card · January 2025
-
Meta model card · December 2024
MMMU (val)
higher is better- Qwen2.5-VL 7B 58.6
Qwen2.5-VL technical report · February 2025
- Llama 3.2 11B Vision 50.7
Meta Llama 3.2 Vision model card (CoT, micro-avg) · September 2024
MTEB Multilingual Mean (Task)
higher is better- Qwen3-Embedding 0.6B 64.33
Qwen3-Embedding-0.6B model card · June 2025
- BGE-M3 59.56
Qwen3-Embedding-0.6B model card (comparison table, reported by Qwen) · June 2025
MTEB-R
higher is better- Qwen3-Reranker 0.6B 65.80
Qwen3-Reranker model card · June 2025
- BGE-reranker-v2-m3 57.03
Qwen3-Reranker card comparison table (competitor-measured) · June 2025
Coding
SWE-bench Verified
higher is better- Muse Glimmer 30B 76.0%
Meta model card · August 2026
- gpt-oss-120b 62.4%
OpenAI gpt-oss model card (high reasoning) · August 2025
- gpt-oss-20b 60.7%
OpenAI gpt-oss model card (high reasoning) · August 2025
- Devstral Small 53.6%
Mistral AI model card (Devstral Small 1.1) · July 2025
- Qwen3-Coder 30B-A3B 51.6
Qwen (official repo, OpenHands scaffold, 100 turns) · August 2025
- Nemotron 3.5 Lightning 51.56
NVIDIA (BF16) · August 2026
Terminal-Bench 2.1
higher is better- Kimi K3 80.9%
Vals leaderboard · August 2026
-
Artificial Analysis (independent) · August 2026
Reasoning
AIME 2024
higher is better- gpt-oss-120b 95.8
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 92.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 81.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 80.4
Qwen3 technical report (thinking mode) · May 2025
- DeepSeek-R1 79.8%
DeepSeek model card · January 2025
- QwQ 32B 79.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 79.3
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 8B 76.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek model card · January 2025
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
MATH-500
higher is better- DeepSeek-R1 97.3%
DeepSeek model card · January 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
MMLU-Pro
higher is better- DeepSeek-R1 84.0
DeepSeek-R1 technical report · January 2025
- Nemotron 3.5 Lightning 81.94
NVIDIA (BF16) · August 2026
- QwQ 32B 69.07
Qwen model card · March 2025
- Gemma 3 27B 67.5
Gemma 3 technical report (27B IT) · March 2025
Rankings are curated and human-approved. Our pipeline can gather current benchmark figures for review, but nothing is published here without a source and a date.