Gemma 3 27B
Google · Text generation · 27B · 131k context · Released 12 March 2025
Gemma 3 27B is the sweet spot of Google's open model family: capable enough to be a genuine general-purpose model, small enough to run on one high-end consumer card. It accepts images alongside text and has a 128k context window. Google publishes quantisation-aware-trained checkpoints, so the 4-bit versions hold up better than a naive quantisation would.
Strengths
- Strong general model that also understands images
- Fits a single 24GB card at 4-bit, with official quantisation-aware weights
- Very broad language support, over 140 languages
Weaknesses
- The Gemma licence carries use conditions, unlike Apache 2.0 models
- Image understanding is capable but not its main strength
- A 27B model needs a good card, not a laptop, for usable quality
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~15GB | Fits a 24GB card, official quantisation-aware weights available |
| Q5_K_M | ~19GB | Better quality, still fits 24GB with modest context |
| Q8_0 | ~28GB | Near-lossless, needs 32GB or more |
| FP16 | ~54GB | Full precision, server or multi-GPU territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
Q4_K_M · ~15GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →For good quality
Q5_K_M · ~19GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →Best quality
FP16 · ~54GB needed
96GB of unified memory
AMD Ryzen AI Max+ 395 (Strix Halo)or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .
Unified memory is shared with the model, so it is already counted above.
Licence
Gemma Terms of Use — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| GPQA Diamond | 42.4 | Gemma 3 technical report (27B IT) | March 2025 |
| MMLU-Pro | 67.5 | Gemma 3 technical report (27B IT) | March 2025 |
| Artificial Analysis Intelligence Index v4.1.1 | 7 | Artificial Analysis (independent) | August 2026 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
Artificial Analysis Intelligence Index v4.1.1
higher is better- gpt-oss-120b 24
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent) · August 2026
- Qwen3 32B 11
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent, figure marked estimated) · August 2026
-
Artificial Analysis (independent) · August 2026
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
MMLU-Pro
higher is better- DeepSeek-R1 84.0
DeepSeek-R1 technical report · January 2025
- Nemotron 3.5 Lightning 81.94
NVIDIA (BF16) · August 2026
- QwQ 32B 69.07
Qwen model card · March 2025
- Gemma 3 27B 67.5
Gemma 3 technical report (27B IT) · March 2025
Gemma 3 27B: common questions
- What hardware do I need to run Gemma 3 27B?
- At its most compressed (Q4_K_M) it needs roughly 15GB of VRAM, and about 19GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is Gemma 3 27B free for commercial use?
- Commercial use is permitted, but with conditions. Gemma is free to use commercially, but Google's Gemma Terms of Use include a Prohibited Use Policy, and you must pass the same use restrictions on to anyone you share the model or a fine-tune with. Read the licence before relying on it at scale.
- Can I run Gemma 3 27B on Apple Silicon?
- Yes. Gemma 3 27B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- Does Gemma 3 27B run on CPU?
- Yes, Gemma 3 27B can run on the CPU, though generation is slower than on a GPU.
- What is Gemma 3 27B's context window?
- Gemma 3 27B has a context window of 131,072 tokens, about 131k.
Availability
- Official page
- Hugging Face
- ollama run gemma3:27b
Where to get quantised weights
Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:
- Unsloth GGUF (Dynamic 2.0, imatrix)
Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.
- Bartowski GGUF Q2-Q8 (imatrix)
A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.
Very broad coverage of models, in both static and imatrix GGUF.
- MLX community MLX 4-bit and 8-bit
MLX quants for Apple Silicon, usually 4-bit and 8-bit.
Recommended for
- A single-card 24GB general model that also handles images
- Multilingual work across many languages
- Users who want a capable Google model with official 4-bit weights
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.