Skip to content
local-ai
Enthusiast

48GB workstation: one card for 70B-class models

A single 48GB professional card runs 70B-class models on one GPU, with no multi-card complexity. It is expensive, but for a quiet, reliable workstation that holds large models, it is the cleanest option short of a server.

Key hardware

Between a 32GB consumer flagship and a data-centre server sits a useful middle ground: a single 48GB professional card. It runs models that a 24GB or 32GB card cannot, without the power, heat, and configuration overhead of a multi-GPU rig.

Who it is for

People who want to run 70B-class models on one card, reliably and quietly, and who would rather pay a professional-card premium than manage two consumer GPUs. It is the step up from the single-card flagship build when 32GB is not enough.

The key part

The RTX A6000 (Ampere) and the newer RTX 6000 Ada both bring 48GB of ECC memory on a single card at a modest 300W. That is enough to run a 70B model at a 4-bit quantisation without splitting it across GPUs. The A6000 is older and slower but often far cheaper on the used market; the Ada card adds FP8 and more bandwidth at a higher price.

What it runs

  • Llama 3.3 70B at Q4 or Q5 with real context headroom
  • gpt-oss-120b is out of reach on 48GB; that needs 80GB
  • Any smaller model at high quantisation with long context

The trade-offs

The cost per gigabyte is far worse than a consumer card, and 48GB still does not reach the 80GB needed for the largest open models. But for a single-card machine that holds 70B-class models quietly and reliably, without multi-GPU complexity, it is the clean answer. If you need more, the next step is a data-centre card or a rented server-class setup.

Build last reviewed 18 August 2026.