Gemma 4 31B
Google · United States · Apache 2.0
Commercial use: Yes — free for commercial use
Apache 2.0 — and this is a change worth knowing about. Gemma 3 shipped under Google's custom Gemma Terms of Use behind a manual access gate. Gemma 4 is tagged apache-2.0 on the Hub and is ungated. The remote-restriction clause and the flow-down obligation that made older Gemma unacceptable to some legal teams do not apply here. Verified against the Hub on 2026-09-06; if you are working from guidance written before Gemma 4, it is out of date.
Google's open model family, and as of Gemma 4 it is Apache 2.0 rather than a custom Google licence.
Why it should matter to you
Licences move in both directions, and this one moved in your favour. Gemma 3 carried terms that let Google restrict use remotely and had to be passed to anyone you redistributed to, which was enough for some legal teams to refuse it outright. Gemma 4 dropped that. If your organisation rejected Gemma on licensing grounds, that decision was made against terms that no longer apply.
How it connects to our work
We re-read licences rather than carrying forward last year's answer, because the family name is not the licence. In the same week we found Gemma loosening, GLM-5.3 tightened — the flagship left MIT while its Flash sibling kept it, and Qwen's largest models left Apache 2.0 while the 27B kept it.
Advantages
- Apache 2.0 and ungated, where Gemma 3 was custom-licensed and required requesting access
- Excellent quality per parameter among Western open models
- Vision capable, 140+ languages
- Free tier on OpenRouter for both the 31B and the 26B-A4B MoE variant
Disadvantages
- Context quality degrades well before the 262K headline figure
- Smaller fine-tune ecosystem than Qwen
- Google has deprecated open model lines before, so plan a fallback
Reach for it when
Multilingual and vision work on a single card, now with a licence that clears OSI-style legal review. One of the strongest Western open options at this size.
Where it falls down
Long-context work at anything near the advertised 262K — test at your real document lengths. Also do not assume older Gemma guidance applies: the licence changed under the same family name.
Running it
24 GB GPU at Q4_K_M, or a 32 GB Mac. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.
Free tiers carrying this model
| Provider | Type | The catch | Live limits |
|---|---|---|---|
| OpenRouter | Aggregator | Free-pool membership changes without notice — a model you built on can stop being free, and the list above will drift. Data passes through OpenRouter and then the upstream provider, so check the privacy and routing settings before sending anything sensitive. | check → |
| NVIDIA NIM | Inference host | It is a credit grant, not a permanent free tier. When the credits run out you are on a paid plan or self-hosting the container. | check → |
| Cloudflare Workers AI | Edge inference | The Neuron accounting makes it hard to predict how many actual requests you get, because different models consume at very different rates. | check → |
| Google AI Studio / Gemini API | First-party API | This is the important one: on the free tier Google may use your prompts and responses to improve its products. Paid tiers do not. If you are sending anything confidential, the free tier is the wrong place for it. | check → |
Jurisdiction: United States
Best raw capability and the deepest tooling ecosystem. For EU personal data you are relying on a transfer framework rather than on data never leaving the bloc, so check whether your DPA and your customers accept that.
Watch for: Enterprise API tiers usually promise no training on your data; consumer tiers and free tiers frequently do not. The free tier is where this bites.
Frequently asked
Can I use Gemma 4 31B commercially?
Apache 2.0 — and this is a change worth knowing about. Gemma 3 shipped under Google's custom Gemma Terms of Use behind a manual access gate. Gemma 4 is tagged apache-2.0 on the Hub and is ungated. The remote-restriction clause and the flow-down obligation that made older Gemma unacceptable to some legal teams do not apply here. Verified against the Hub on 2026-09-06; if you are working from guidance written before Gemma 4, it is out of date.
What hardware do I need to run Gemma 4 31B?
24 GB GPU at Q4_K_M, or a 32 GB Mac. Weights alone are roughly 18 GB at Q4_K_M and 33 GB at Q8_0. Add KV cache on top of that, which grows with your context length.
What licence is Gemma 4 31B released under?
Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.
Where can I use Gemma 4 31B for free?
Free tiers carrying it include OpenRouter, NVIDIA NIM, Cloudflare Workers AI, Google AI Studio / Gemini API. Limits differ per provider and change often, so check each provider's own limits page. You can also self-host the weights, which has no rate limit at all.
Similar models
Fara-7B
8 GB GPU or any 16 GB Mac.
Llama 4 Scout 17B
80 GB GPU or 96 GB+ Mac at Q4_K_M.
gpt-oss-120b
80 GB GPU, or 96 GB+ unified memory.
gpt-oss-20b
16 GB GPU or Mac. A strong default for local agents.
Wiring Gemma 4 31B into something real?
We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.
Maps to AI agent development, AI automation development and AI strategy and consulting. Or see it working: our case studies.