local-ai.net
Run AI on your own hardware.
News, guides, and a catalogue of models, tools, and hardware for local AI. Plus two tools that answer the questions people actually arrive with: what you can run, and whether running it locally is cheaper.
Hardware capability matrix
Tell it your GPU or memory, and see which models fit and at what quantisation. Figures are approximate; fits does not always mean runs well.
Local vs API cost calculator
Work out the break-even point between buying hardware and paying for an API, with the assumptions stated plainly, including the ones that count against local.
Latest news
All news-
6 July 2026
Mistral releases the Mistral 3 family under Apache 2.0
Mistral has released the Mistral 3 family, spanning small dense models and a large mixture-of-experts, all under the Apache 2.0 licence. The licensing is as notable as the models: it removes the commercial conditions of Mistral's earlier releases.
-
2 July 2026
Mistral open-sources Leanstral, a model built for formal verification
Mistral has released Leanstral 1.5, an Apache 2.0 model specialised for Lean 4, the proof assistant used in formal software verification. It is a reminder that local models are not only general chatbots, but increasingly tools for specific domains.
-
29 June 2026
EU finalises Digital Omnibus, delaying the AI Act's high-risk rules to 2027
The Council of the EU has given final approval to the Digital Omnibus, a package that simplifies and defers parts of the AI Act. The main high-risk obligations move from August 2026 to December 2027, while the prohibitions and model rules stay live.
-
18 May 2026
Apple trims high-memory Mac Studio options amid a DRAM squeeze
A tightening memory market has led Apple to narrow the memory configurations on its top Mac Studio, removing some of the very high-capacity options that made it a favourite for running large models. If you were counting on a specific config, check current availability.
Featured models
Full catalogue-
Gemma 3 27B
Google · 27B
Google's 27-billion-parameter open model, and the largest of the Gemma 3 family. It is multimodal, handling images as well as text, supports over 140 languages, and is designed to run on a single 24GB card at 4-bit.
Permitted with conditions -
Llama 3.3 70B Instruct
Meta · 70B
Meta's 70-billion-parameter instruction-tuned model, delivering performance close to their much larger 405B model at a fraction of the hardware cost. A strong general-purpose choice if you have the VRAM for it.
Permitted with conditions -
Qwen2.5-Coder 7B Instruct
Alibaba · 7B
Alibaba's code-specialised 7B model, strong at code completion and generation well beyond what its size would suggest. A practical choice for a local coding assistant on a mid-range GPU, and permissively licensed.
Commercial use permitted -
Qwen3 30B-A3B
Alibaba · MoE 30B-A3B
A mixture-of-experts model with 30 billion total parameters but only 3 billion active per token, so it generates far faster than a dense 30B while keeping much of the quality. Apache 2.0, with an optional thinking mode.
Commercial use permitted -
Qwen3 32B
Alibaba · 32B
The largest dense Qwen3 model, and a strong single-card flagship: at Q4 it fits a 24GB card with short context. Apache 2.0 licensed, with an optional thinking mode that makes it a credible reasoning model.
Commercial use permitted -
Qwen3 8B
Alibaba · 8B
Alibaba's 8-billion-parameter model from the Qwen3 family, with a hybrid thinking mode you can switch on for harder problems. Apache 2.0 licensed and strong for its size, it is one of the most capable models that runs comfortably on a 12GB card.
Commercial use permitted -
Qwen3-Coder 30B-A3B
Alibaba · MoE 30B-A3B
A code-specialised mixture-of-experts model with 30 billion total parameters and 3 billion active, built for agentic coding: exploring codebases, editing across files, and driving coding agents. Apache 2.0, with a very long 256k context window.
Commercial use permitted
Learning paths
- Getting Started What local AI is, why it matters, and how to run your first model.
- Hardware GPUs, unified memory, VRAM, and what you actually need for what you want to do.
- Running Models Quantisation, inference engines, model formats, and getting good performance.
- Local AI Coding Code assistants, autocomplete, and agentic coding tools that run on your machine.
- RAG & Knowledge Systems Building retrieval systems over your own documents, entirely locally.
- Agentic AI & Harnesses Tool use, agent frameworks, and orchestration with local models.
- Fine-Tuning Adapting models to your domain: LoRA, QLoRA, full fine-tunes, and when each is worth it.
- Beyond LLMs Image generation, speech recognition, text-to-speech, and other local models.
- Server & Enterprise Running local AI at organisational scale: serving, scaling, and multi-user access.
- AI Regulation & Sovereignty AI-specific regulation like the EU AI Act, alongside data residency, GDPR, and the practical case for keeping data in-house.
Local AI is built by people who think running AI on your own hardware matters, for privacy, autonomy, and an open ecosystem. That belief shapes what we cover, not how we report it. We are honest about where local AI falls short and where cloud is the better answer.