Skip to content
local-ai

Gear

What to actually buy to run AI on your own hardware, by scenario. Every recommendation is made on merit; where a link earns us a commission it is marked, and it never changes what we recommend. When the honest answer is "rent, do not buy", we say so.

Work it out for yourself

Recommendations by scenario

Prices move constantly, so we stay on tiers and let you check current retail through the links. VRAM figures are approximate.

Front-end developer: Qwen3.8-27B at 40+ tokens/second

A capable coding and reasoning assistant, fast enough to feel instant.

24GB is the sweet spot for a 27B model at a 4-bit quantisation. For value, a used RTX 3090 (24GB) is hard to beat; for the most speed and newest architecture, the RTX 5090 (32GB). To comfortably clear 40 tokens per second on a 27B, a fast discrete GPU is a surer bet than unified memory, which trades speed for capacity. See Qwen3.8-27B and check the numbers on the machine speccer. Builds: best-value 24GB or single-card flagship.

Shop the RTX 5090 on Amazon (affiliate)

Small business: running Kimi K3 in-house for documents and software engineering

You want a top open model on your own infrastructure for mixed knowledge work and complex coding.

Here the honest answer matters more than a buy link. Kimi K3 is a 2.8-trillion-parameter model: genuinely server-class, needing a multi-node cluster rather than a workstation. For most small businesses the sensible route is to rent server GPUs by the hour rather than buy a cluster that sits idle. Run the numbers on the cost calculator, and read running large open models.

If a somewhat smaller model will do, you gain a lot of practicality: gpt-oss-120b runs on a single 80GB data-centre card and is far easier to self-host, and a strong 32B like Qwen3 32B serves a small team well from one 24GB to 48GB card. That is usually the better first move than a Kimi-scale cluster.