Skip to content
local-ai

NVIDIA Tesla P40

NVIDIA · Budget

A 2016 Pascal data-centre accelerator with 24GB of memory, repurposed as a cheap high-VRAM card for homelab local inference.

Specifications

VRAM 24GB
Memory bandwidth 346 GB/s
Power draw 250W
Type accelerator

A cheap used homelab card, though cooling and power-adapter costs add up. Verify current pricing.

NVIDIA Tesla P40: common questions

What models can the NVIDIA Tesla P40 run?
With 24GB of VRAM it can run models up to roughly 37B parameters at a 4-bit quantisation, or smaller models with more context. Use the hardware matrix for specifics; these figures are approximate.
How much power does the NVIDIA Tesla P40 draw?
About 250W under load, so pair it with a power supply that has real headroom.

Strengths

  • 24GB at a very low used price, the classic budget entry to larger local models
  • Widely supported in llama.cpp and other inference stacks
  • A passive 250W design suits a server chassis with existing airflow

Weaknesses

  • Only 346 GB/s bandwidth, so generation is slow versus modern cards
  • Weak FP16 and no tensor cores, so it suits quantised GGUF workloads
  • The passive cooler needs added fans in a desktop, and it uses a server power connector
See what models fit in 24GB →

Entry last verified 18 August 2026. Specifications and especially prices change; verify before buying.