Skip to content
local-ai

NVIDIA H200

NVIDIA · Server

A Hopper-architecture data-centre GPU, the first with 141GB of HBM3e at 4.8 TB/s, built for large-model training and high-throughput serving.

Specifications

VRAM 141GB
Memory bandwidth 4800 GB/s
Power draw 700W
Type accelerator

An enterprise GPU, typically rented by the hour rather than bought. Verify current pricing.

NVIDIA H200: common questions

What models can the NVIDIA H200 run?
With 141GB of VRAM it can run models up to roughly 232B parameters at a 4-bit quantisation, or smaller models with more context. Use the hardware matrix for specifics; these figures are approximate.
How much power does the NVIDIA H200 draw?
About 700W under load, so pair it with a power supply that has real headroom.

Strengths

  • 141GB holds very large models on one GPU, with room for long context and KV cache
  • 4.8 TB/s bandwidth delivers throughput no consumer card approaches
  • FP8 and NVLink for multi-GPU scaling in server nodes

Weaknesses

  • Data-centre only: the SXM form factor needs an HGX baseboard, not a desktop
  • A 700W draw demands server-grade power and cooling
  • Cost puts it far outside individual budgets, so it is realistically rented
See what models fit in 141GB →

Entry last verified 18 August 2026. Specifications and especially prices change; verify before buying.