Model catalogue
Models you can run on your own hardware, from a laptop to a server rack, with plain-English licence classifications, approximate memory requirements, and honest strengths and weaknesses. Filter by type, size, modality, licence, developer, and format, or sort by release date to see the latest first.
45 models.
Z.ai (Zhipu) · Vision-language · MoE 320B (A18B) · Released Aug 2026
The first natively multimodal model in the GLM-5 line: a roughly 320-billion-parameter mixture-of-experts with about 18 billion active per token, handling text, images, video and visual documents. MIT-licensed and around half the size of the full GLM-5.3, though still server-class.
Commercial use permittedText + Image + VideoTextfrom ~320GBVery largeZ.ai (Zhipu) · Text generation · MoE 744B (A40B) · Released Aug 2026
Z.ai's server-class, coding-focused reasoning model: a roughly 744-billion-parameter mixture-of-experts with about 40 billion active per token and a 1-million-token context. Open-weight under a bespoke GLM-5.3 licence rather than MIT, and firmly cluster-scale rather than a consumer-card model.
Permitted with conditionsfrom ~750GBVery largeAlibaba · Text generation · 27B · Released Aug 2026
The runnable member of the Qwen3.8 line: a 27-billion-parameter dense model under Apache 2.0, natively multimodal with image and video input, and a 262k context window that extends towards a million tokens with YaRN. It fits a 24GB card at Q4, which makes it a strong local all-rounder.
Commercial use permittedText + Image + VideoTextfrom ~17GBMediumGGUFMLXAlibaba · Text generation · MoE 2.4T · Released Aug 2026
The open-weight release of Alibaba's flagship, a 2.4-trillion-parameter mixture-of-experts with 95 billion active per token. It is firmly server-class, needing a cluster rather than a workstation, and it ships under a custom licence rather than the Apache 2.0 used across the rest of the Qwen line.
Permitted with conditionsfrom ~4800GBVery largeLiquid AI · Vision-language · 3.1B · Released Aug 2026
A 3.1-billion-parameter open-weight vision-language model from Liquid AI, built to run on-device: phones, laptops, and single consumer GPUs rather than in a data centre. It adds screen understanding, object grounding, multi-image reasoning, and function calling, and ships with day-one GGUF, MLX, and ONNX exports.
Permitted with conditionsText + ImageTextfrom ~3GBTinyGGUFMLXNVIDIA · Text generation · MoE 31.6B-A3.6B · Released Aug 2026
NVIDIA's fast, fully open mixture-of-experts: about 31.6 billion parameters with 3.6 billion active per token, a hybrid Mamba-Transformer design, and a one-million-token context. Released under a permissive licence with the training data and recipes alongside the weights, and aimed at high-volume agentic work.
Commercial use permittedfrom ~18GBMediumMeta · Text generation · 30B · Released Aug 2026
Meta's 30-billion-parameter open model, distilled from its larger closed Muse Spark and built for agentic work: tool use, multi-step reasoning, and recovering from failures. It accepts images as well as text, runs on a single 24GB card at 4-bit, and, unusually for Meta, ships under Apache 2.0 rather than the Llama licence.
Commercial use permittedText + ImageTextfrom ~18GBMediumGGUFMoonshot AI · Text generation · MoE 2.8T · Released Jul 2026
Moonshot AI's flagship open-weight model, a 2.8-trillion-parameter mixture-of-experts that ranks at or near the top of independent open-weight leaderboards. It is firmly server-class: running it needs a cluster, not a workstation, and its licence terms were not clearly published at launch.
Licence unclearfrom ~1400GBVery largeOpenAI · Text generation · MoE 117B-A5B · Released Aug 2025
OpenAI's larger open-weight model, a mixture-of-experts with about 117 billion total parameters but only 5.1 billion active per token. Apache 2.0, and notable for running on a single 80GB GPU thanks to a native 4-bit format, while OpenAI reports reasoning near its o4-mini.
Commercial use permittedfrom ~61GBLargeGGUFMLXOpenAI · Text generation · MoE 21B-A3.6B · Released Aug 2025
The smaller of OpenAI's open-weight models, a mixture-of-experts with about 21 billion total parameters and 3.6 billion active. Apache 2.0, and it runs in roughly 16GB thanks to a native 4-bit format, with reasoning OpenAI compares to its o3-mini.
Commercial use permittedfrom ~13GBSmallGGUFMLXAlibaba · Image generation · 20B · Released Aug 2025
Alibaba's 20-billion-parameter image model, notable for strong text rendering and a permissive Apache 2.0 licence. It is designed to fit under 16GB with quantisation, and generates a 1024px image in a handful of steps.
Commercial use permittedTextImagefrom ~14GBSmallGGUFMLXAlibaba · Video · 5B · Released Jul 2025
The consumer-runnable member of Alibaba's Wan 2.2 video family: a 5-billion-parameter model that does both text-to-video and image-to-video at 720p, and, unusually for open video, fits a single 24GB card. Apache 2.0.
Commercial use permittedText + ImageVideofrom ~24GBMediumMistral AI · Code · 24B · Released Jul 2025
A 24-billion-parameter coding model from Mistral AI and All Hands AI, built specifically to drive software engineering agents like OpenHands. Apache 2.0, and designed to run on a single 24GB card or a 32GB Mac.
Commercial use permittedfrom ~14GBSmallGGUFMLXAlibaba · Embedding · 0.6B · Released Jun 2025
The smallest model in Alibaba's Qwen3-Embedding family, strong on multilingual retrieval for its size while staying light enough to run on a CPU. Apache 2.0, with user-selectable output dimensions.
Commercial use permittedfrom ~0.7GBTinyGGUFMLXAlibaba · Reranker · 0.6B · Released Jun 2025
A small cross-encoder reranker from the Qwen3 Embedding series that re-scores retrieved passages by how well each actually answers a query, the quality-boosting second stage of a RAG pipeline. Apache 2.0, multilingual, and tiny enough to run almost anywhere.
Commercial use permittedfrom ~2GBTinyMistral AI · Text generation · 24B · Released Jun 2025
Mistral's strong 24B general-purpose model, and a rare thing in the catalogue: a capable, permissively-licensed European alternative to the Qwen and Llama families. Apache 2.0, multimodal, with a 128k context and a refined 3.2 update that improves instruction following and reliability.
Commercial use permittedText + ImageTextfrom ~15GBSmallGGUFMLXAlibaba · Code · MoE 30B-A3B · Released May 2025
A code-specialised mixture-of-experts model with 30 billion total parameters and 3 billion active, built for agentic coding: exploring codebases, editing across files, and driving coding agents. Apache 2.0, with a very long 256k context window.
Commercial use permittedfrom ~18GBMediumGGUFMLXAlibaba · Text generation · 14B · Released Apr 2025
The 14-billion-parameter Qwen3 model, a strong middle ground: noticeably more capable than the 8B while still fitting a 16GB card at a good quantisation. Apache 2.0, with an optional thinking mode and 128k context.
Commercial use permittedfrom ~9GBSmallGGUFMLXAlibaba · Text generation · MoE 30B-A3B · Released Apr 2025
A mixture-of-experts model with 30 billion total parameters but only 3 billion active per token, so it generates far faster than a dense 30B while keeping much of the quality. Apache 2.0, with an optional thinking mode.
Commercial use permittedfrom ~18GBMediumGGUFMLXAlibaba · Text generation · 32B · Released Apr 2025
The largest dense Qwen3 model, and a strong single-card flagship: at Q4 it fits a 24GB card with short context. Apache 2.0 licensed, with an optional thinking mode that makes it a credible reasoning model.
Commercial use permittedfrom ~20GBMediumGGUFMLXAlibaba · Text generation · 4B · Released Apr 2025
A 4-billion-parameter model from the Qwen3 family that punches above its weight, with an optional thinking mode and a 128k context window. Light enough for laptops and 8GB cards, and Apache 2.0 licensed.
Commercial use permittedfrom ~2.5GBTinyGGUFMLXAlibaba · Text generation · 8B · Released Apr 2025
Alibaba's 8-billion-parameter model from the Qwen3 family, with a hybrid thinking mode you can switch on for harder problems. Apache 2.0 licensed and strong for its size, it is one of the most capable models that runs comfortably on a 12GB card.
Commercial use permittedfrom ~5GBTinyGGUFMLXGoogle · Text generation · 27B · Released Mar 2025
Google's 27-billion-parameter open model, and the largest of the Gemma 3 family. It is multimodal, handling images as well as text, supports over 140 languages, and is designed to run on a single 24GB card at 4-bit.
Permitted with conditionsText + ImageTextfrom ~15GBSmallGGUFMLXAlibaba · Text generation · 32B · Released Mar 2025
Alibaba's dedicated reasoning model at 32 billion parameters. It works through problems with extended chain-of-thought before answering, reaching results competitive with much larger reasoning models. Apache 2.0, and runnable on a 24GB card.
Commercial use permittedfrom ~20GBMediumGGUFMLXAlibaba · Vision-language · 7B · Released Jan 2025
A 7-billion-parameter vision-language model that reads images and documents well above its weight, with particularly strong optical character recognition and chart and table understanding. Apache 2.0, and runnable on a 12GB card.
Commercial use permittedText + ImageTextfrom ~6GBTinyGGUFMLXDeepSeek · Text generation · MoE 671B-A37B · Released Jan 2025
DeepSeek's flagship open-weight reasoning model: a 671-billion-parameter mixture-of-experts with 37 billion active per token, reaching results competitive with the strongest proprietary reasoning models. Fully open under the MIT licence, but firmly server territory to run.
Commercial use permittedfrom ~140GBVery largeGGUFMLXDeepSeek · Text generation · 32B · Released Jan 2025
A reasoning model distilled from DeepSeek-R1 into a 32B Qwen2.5 base, bringing much of R1's chain-of-thought reasoning to hardware that can run a 32B model. A practical way to get strong local reasoning on a 24GB card.
Commercial use permittedfrom ~20GBMediumGGUFMLXHexgrad · Text to speech · 82M · Released Jan 2025
A tiny 82-million-parameter text-to-speech model that produces natural 24kHz speech and runs fast on a laptop CPU. Apache 2.0 licensed, and in blind tests it holds its own against models many times its size.
Commercial use permittedTextAudiofrom ~0.3GBTinyMeta · Text generation · 70B · Released Dec 2024
Meta's 70-billion-parameter instruction-tuned model, delivering performance close to their much larger 405B model at a fraction of the hardware cost. A strong general-purpose choice if you have the VRAM for it.
Permitted with conditionsfrom ~43GBLargeGGUFMLXLightricks · Video · 2B · Released Nov 2024
Lightricks' fast, DiT-based video model, designed for near-real-time generation. The 2B variant is aimed at light VRAM use, making it one of the more approachable open video models, though its per-version weights licence needs reading before commercial use.
Permitted with conditionsText + ImageVideofrom ~12GBSmallAlibaba · Code · 32B · Released Nov 2024
The flagship of the Qwen2.5-Coder family, and for a long time the strongest open coding model you could run on a single 24GB card. Apache 2.0, with a 128k context window and strong scores across code generation and editing benchmarks.
Commercial use permittedfrom ~20GBMediumGGUFMLXAlibaba · Code · 7B · Released Nov 2024
Alibaba's code-specialised 7B model, strong at code completion and generation well beyond what its size would suggest. A practical choice for a local coding assistant on a mid-range GPU, and permissively licensed.
Commercial use permittedfrom ~4.7GBTinyGGUFMLXMeta · Vision-language · 11B · Released Sep 2024
Meta's 11-billion-parameter vision-language model, a solid general choice for understanding images alongside text on a mid-range card. Well supported across local tooling, with a long 128k context window.
Permitted with conditionsText + ImageTextfrom ~8GBSmallGGUFMLXMeta · Text generation · 3B · Released Sep 2024
A small 3-billion-parameter model that runs on almost anything, including laptops without a dedicated GPU. Capable for its size at summarising, rewriting, and simple assistant tasks, and a good first model to try.
Permitted with conditionsfrom ~2.2GBTinyGGUFMLXBlack Forest Labs · Image generation · 12B · Released Aug 2024
The higher-quality member of the FLUX.1 family, producing excellent photorealistic images with readable text and strong prompt adherence. The catch is licensing: it is released under a non-commercial licence, so commercial use needs a separate agreement with Black Forest Labs.
Non-commercial onlyTextImagefrom ~8GBSmallGGUFMLXBlack Forest Labs · Image generation · 12B · Released Aug 2024
The fast, Apache 2.0 member of Black Forest Labs' FLUX.1 family. It generates high-quality images, including genuinely readable text, in as few as one to four steps, which keeps it quick even on mid-range hardware.
Commercial use permittedTextImagefrom ~8GBSmallGGUFMLXStability AI · Audio & music · 1B · Released Jul 2024
Stability AI's open text-to-audio model, generating up to 47 seconds of stereo 44.1kHz audio. It excels at sound effects and field recordings rather than vocals, and is free for research and for smaller organisations, with an enterprise licence above a revenue threshold.
Permitted with conditionsTextAudiofrom ~6GBTinyNomic AI · Embedding · 137M · Released Feb 2024
A small, fast English-focused embedding model with an 8192-token input, and one of the most common defaults for local RAG. Apache 2.0 and light enough to run comfortably on a CPU.
Commercial use permittedfrom ~0.3GBTinyGGUFBAAI · Reranker · 0.6B · Released Feb 2024
A widely-used, lightweight multilingual reranker from BAAI, built on the BGE-M3 base. It re-scores retrieved passages by relevance to sharpen a RAG pipeline, and its long track record makes it a safe, well-supported default.
Commercial use permittedfrom ~2GBTinyBAAI · Embedding · 568M · Released Jan 2024
A multilingual embedding model that is unusual for doing dense, sparse, and multi-vector retrieval from a single model, across more than 100 languages, with an 8192-token input. MIT licensed and a common default for local RAG.
Commercial use permittedfrom ~0.7GBTinyGGUFOpenAI · Speech to text · 1.55B · Released Nov 2023
OpenAI's flagship open speech-to-text model, and the de facto standard for local transcription. It handles many languages, is robust to accents and background noise, and runs well through whisper.cpp or faster-whisper. MIT licensed.
Commercial use permittedAudioTextfrom ~1.5GBTinyGGUFMLXStability AI · Image generation · 3.5B · Released Jul 2023
The workhorse of local image generation. SDXL produces native 1024x1024 images and has by far the deepest community library of fine-tunes and LoRAs of any base model, so it is endlessly customisable. It runs on 8GB and up.
Permitted with conditionsTextImagefrom ~7GBTinyMLXACE Studio and StepFun · Audio & music · 3.5B
An open, Apache 2.0 music generation model that produces full songs with duration control, up to around four minutes. It runs in as little as 8GB of VRAM, which, combined with its permissive licence, makes it a practical local choice for music.
Commercial use permittedTextAudiofrom ~8GBSmallTencent · Video · 13B
Tencent's 13-billion-parameter open video model, producing text-to-video up to 720p. It is a high-quality but high-VRAM option, needing 45GB or more, and ships under a custom community licence rather than a standard open-source one.
Commercially restrictedTextVideofrom ~35GBLargeMeta · Audio & music · 3.3B
Meta's text-to-music model, available in 300M, 1.5B, and 3.3B sizes plus a melody-guided variant. The important catch is licensing: the weights are released for non-commercial use only, so it cannot be used in commercial work.
Non-commercial onlyTextAudiofrom ~8GBSmall