Single-card flagship: fast inference on 32GB
A no-compromises single-card build around the RTX 5090. Its 32GB and high bandwidth run strong models up to around 32B quickly, with room for longer context than a 24GB card allows. Fast, capable, and expensive.
Key hardware
- NVIDIA GeForce RTX 5090 — 32GB
This build is for people who want the best single-card experience: fast generation, and enough memory to run strong models without dropping to aggressive quantisations. It is not about value, it is about capability on one card.
Who it is for
Enthusiasts and professionals who use local models heavily and want them to feel quick, and who can hold longer context than a 24GB card comfortably allows. If you want the speed and do not want to manage a multi-GPU setup, this is the build.
The key part
The RTX 5090 brings 32GB of fast GDDR7 and very high memory bandwidth. The extra memory over a 24GB card matters in two ways: it fits slightly larger models, and, more usefully, it leaves real headroom for long context on models in the 32B class. The high bandwidth is what makes generation feel fast.
Give it a capable modern CPU, 64GB of system RAM if you can, and a power supply with plenty of headroom, as the card draws a lot under load and runs hot. Cooling deserves real thought.
What it runs
- Qwen3 32B at a good quantisation, with usable context
- Qwen3 30B-A3B very quickly, thanks to the bandwidth
- Larger image models like FLUX.1 [schnell] at speed
The trade-offs
The price is the trade-off, and it is steep. If value matters more than speed, a used 3090 gives the same 24GB tier for far less. If you need to run models beyond 32GB, you are into multi-GPU territory or a high-memory machine. But as a single card that does most things quickly, the 5090 is the top of the consumer tree.
Build last reviewed 30 July 2026.