Skip to content
local-ai

Model rankings

A curated summary of well-regarded models by task and size, drawn from external benchmarks. We do not run our own benchmarks; every entry is attributed and dated so you can judge it yourself.

General chat

General-purpose assistants and instruction following.

Size tier Model Draws on Source As of
Medium Qwen3 32B Alibaba An independent aggregate index. The strongest general model that fits a single 24GB card in this set. Artificial Analysis Intelligence Index v4.1.1: 11 Artificial Analysis August 2026
Medium Mistral Small 3.2 24B Mistral AI A releasing-organisation figure. A permissive, multimodal European alternative on a 24GB card. HumanEval Plus (Pass@5) 92.9% Mistral Small 3.2 model card June 2025
Large Llama 3.3 70B Instruct Meta A releasing-organisation figure. Treat as a claim pending independent confirmation. MMLU 86.0 Meta model card December 2024
Large gpt-oss-120b OpenAI An independent aggregate index across nine evaluations. Runs on a single 80GB card. Artificial Analysis Intelligence Index v4.1.1: 24 Artificial Analysis August 2026

Coding

Code generation, completion, and understanding.

Size tier Model Draws on Source As of
Tiny Qwen2.5-Coder 7B Instruct Alibaba A releasing-organisation figure. HumanEval measures single-function generation, not agentic coding. HumanEval 88.4% Qwen2.5-Coder technical report September 2024
Small Devstral Small Mistral AI A releasing-organisation figure. SWE-bench measures agentic, multi-file software tasks. SWE-bench Verified 53.6% Mistral AI model card (Devstral Small 1.1) July 2025
Medium Qwen3-Coder 30B-A3B Alibaba A releasing-organisation figure; reproducibility is scaffold-sensitive. A mixture-of-experts coder that runs in roughly 18GB. SWE-bench Verified 51.6 Qwen (official repo, OpenHands scaffold) August 2025
Medium Qwen2.5-Coder 32B Alibaba A releasing-organisation figure. The strongest local coder that fits a 24GB card. HumanEval 92.7, Aider Pass@2 73.7 Qwen2.5-Coder technical report / blog November 2024
Large gpt-oss-120b OpenAI An independent leaderboard figure. Aider Polyglot measures multi-language code editing. Aider Polyglot 41.8% Aider polyglot leaderboard August 2026
Very large DeepSeek-R1 DeepSeek An independent leaderboard figure. Server-class hardware to run. Aider Polyglot 56.9% Aider polyglot leaderboard August 2026

Reasoning

Maths, logic, and multi-step problem solving.

Size tier Model Draws on Source As of
Tiny Qwen3 8B Alibaba A releasing-organisation figure. Strong reasoning for a model that fits 8GB. GPQA Diamond 62.0, AIME 2024 76.0 Qwen3 technical report (thinking mode) May 2025
Small Qwen3 14B Alibaba A releasing-organisation figure with thinking mode enabled. GPQA Diamond 64.0, AIME 2024 79.3 Qwen3 technical report (thinking mode) May 2025
Medium QwQ 32B Alibaba A releasing-organisation figure. Treat as a claim pending independent confirmation. MMLU-Pro 69.07 Qwen model card March 2025
Medium DeepSeek-R1-Distill-Qwen 32B DeepSeek A releasing-organisation figure. AIME is a hard maths benchmark; strong here does not imply strong everywhere. AIME 2024 72.6% DeepSeek model card January 2025
Large gpt-oss-120b OpenAI A releasing-organisation figure at high reasoning effort. GPQA Diamond 80.1, AIME 2024 95.8 OpenAI gpt-oss model card (high reasoning) August 2025
Very large DeepSeek-R1 DeepSeek A releasing-organisation figure. Server-class hardware to run; most people use a distill. MMLU 90.8, AIME 2024 79.8% DeepSeek model card January 2025

Embedding

Text embeddings for retrieval and RAG.

Size tier Model Draws on Source As of
Tiny Qwen3 Embedding 0.6B Alibaba A releasing-organisation figure. Tiny and fast; the 8B sibling scores higher. MTEB Multilingual Mean (Task) 64.33 Qwen3-Embedding model card June 2025
Small Qwen3-Embedding 8B Alibaba A leaderboard ranking rather than a self-reported score. The lighter 0.6B version of this family is in our catalogue. MTEB Multilingual 70.58 (No.1) Qwen, MTEB multilingual leaderboard June 2025

Reranking

Re-scoring retrieved passages to sharpen RAG results.

Size tier Model Draws on Source As of
Tiny Qwen3-Reranker 0.6B Alibaba A releasing-organisation figure. Larger 4B and 8B siblings score higher. MTEB-R 65.80 Qwen3-Reranker model card June 2025
Tiny BGE-reranker-v2-m3 BAAI A mature, widely-supported default. Figure measured by a competitor, not BAAI. MTEB-R 57.03 Qwen3-Reranker card comparison (competitor-measured) June 2025

Vision-language

Understanding images alongside text.

Size tier Model Draws on Source As of
Tiny Qwen2.5-VL 7B Alibaba A releasing-organisation figure. Especially strong on documents and OCR. DocVQA (test) 95.7, MMMU (val) 58.6 Qwen2.5-VL technical report February 2025
Small Llama 3.2 11B Vision Meta A releasing-organisation figure (MMMU reported with chain-of-thought). MMMU (val) 50.7, DocVQA (test) 88.4 Meta Llama 3.2 Vision model card September 2024

Image generation

Generating images from text prompts.

No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.

Speech

Speech recognition and synthesis.

Size tier Model Draws on Source As of
Tiny Whisper large-v3 OpenAI An independent leaderboard figure; word error rate, so lower is better. Open ASR Leaderboard average WER 7.44 (lower is better) Open ASR Leaderboard October 2025

Head-to-head on shared benchmarks

Where two or more models we cover report the same benchmark, we line the scores up. Most are developer-reported, from the source and date on each row, so testing conditions differ between them. Scores are comparable within a board, not across boards.

General chat

MMLU

higher is better

Coding

Reasoning

Rankings are curated and human-approved. Our pipeline can gather current benchmark figures for review, but nothing is published here without a source and a date.