Skip to content
local-ai

Gemma 3 27B

Google · Text generation · 27B · 131k context · Released 12 March 2025

Permitted with conditions Text + ImageText Open weights Runs on CPU Apple Silicon

Gemma 3 27B is the sweet spot of Google's open model family: capable enough to be a genuine general-purpose model, small enough to run on one high-end consumer card. It accepts images alongside text and has a 128k context window. Google publishes quantisation-aware-trained checkpoints, so the 4-bit versions hold up better than a naive quantisation would.

Strengths

  • Strong general model that also understands images
  • Fits a single 24GB card at 4-bit, with official quantisation-aware weights
  • Very broad language support, over 140 languages

Weaknesses

  • The Gemma licence carries use conditions, unlike Apache 2.0 models
  • Image understanding is capable but not its main strength
  • A 27B model needs a good card, not a laptop, for usable quality

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~15GBFits a 24GB card, official quantisation-aware weights available
Q5_K_M~19GBBetter quality, still fits 24GB with modest context
Q8_0~28GBNear-lossless, needs 32GB or more
FP16~54GBFull precision, server or multi-GPU territory

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~15GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

For good quality

Q5_K_M · ~19GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

Best quality

FP16 · ~54GB needed

96GB of unified memory

AMD Ryzen AI Max+ 395 (Strix Halo)

or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .

Unified memory is shared with the model, so it is already counted above.

Licence

Gemma Terms of Use read the licence

Benchmarks

BenchmarkScoreSourceAs of
GPQA Diamond42.4 Gemma 3 technical report (27B IT) March 2025
MMLU-Pro67.5 Gemma 3 technical report (27B IT) March 2025
Artificial Analysis Intelligence Index v4.1.17 Artificial Analysis (independent) August 2026

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

Gemma 3 27B: common questions

What hardware do I need to run Gemma 3 27B?
At its most compressed (Q4_K_M) it needs roughly 15GB of VRAM, and about 19GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Gemma 3 27B free for commercial use?
Commercial use is permitted, but with conditions. Gemma is free to use commercially, but Google's Gemma Terms of Use include a Prohibited Use Policy, and you must pass the same use restrictions on to anyone you share the model or a fine-tune with. Read the licence before relying on it at scale.
Can I run Gemma 3 27B on Apple Silicon?
Yes. Gemma 3 27B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Gemma 3 27B run on CPU?
Yes, Gemma 3 27B can run on the CPU, though generation is slower than on a GPU.
What is Gemma 3 27B's context window?
Gemma 3 27B has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

  • Very broad coverage of models, in both static and imatrix GGUF.

  • MLX community MLX 4-bit and 8-bit

    MLX quants for Apple Silicon, usually 4-bit and 8-bit.

Recommended for

  • A single-card 24GB general model that also handles images
  • Multilingual work across many languages
  • Users who want a capable Google model with official 4-bit weights

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.