Quiet large-model box: 128GB unified memory at low power
A compact, low-power route to running very large models, built around a Ryzen AI Max+ 395 mini PC with 128GB of unified memory. It holds models no single consumer GPU can, in exchange for slower generation.
Key hardware
This build takes the opposite approach to a discrete GPU. Instead of a fast card with limited memory, it uses a large pool of unified memory to hold very large models, quietly and at low power, accepting slower generation as the price.
Who it is for
People who want to run large models at home without a noisy, power-hungry multi-GPU rig, and who value quiet and efficiency over raw speed. It suits running a big model as a steady local assistant more than it suits fast, interactive bursts.
The key part
The Ryzen AI Max+ 395, codenamed Strix Halo, is sold as a complete mini PC. On the 128GB configuration, up to 96GB can be assigned as memory for models, which is enough to hold models that no single consumer discrete card can fit. It does this at low power, in a small, quiet box.
Because it is a complete machine, there is nothing else to build. The trade you are making is bandwidth: it is well below a discrete GPU, so generation is slower.
What it runs
- Large models such as a 70B at reasonable quantisations
- Qwen3 30B-A3B, whose mixture-of-experts design suits this hardware well, keeping generation responsive
- Comfortable headroom for long context on mid-sized models
The trade-offs
Slower generation is the honest cost, and the wider AMD software stack is less mature than NVIDIA’s CUDA, so expect occasional friction. It occupies similar ground to a high-memory Apple Silicon build, usually at a lower price. If you want speed rather than capacity, a discrete-GPU build serves you better. For quietly running large models at home, this is a genuinely new and compelling option.
Build last reviewed 30 July 2026.