homeservicesworkaboutblogfree templatescontactFree Tools →Free AI ModelsResearch LibraryROI CalculatorSavings CalculatorAI Readiness ScoreHire vs. AutomateAutomation Quote
book a 30-min call
home / free-models / gemma-4-31b

Gemma 4 31B

Google · United States · Apache 2.0

Commercial use: Yes — free for commercial use

Apache 2.0 — and this is a change worth knowing about. Gemma 3 shipped under Google's custom Gemma Terms of Use behind a manual access gate. Gemma 4 is tagged apache-2.0 on the Hub and is ungated. The remote-restriction clause and the flow-down obligation that made older Gemma unacceptable to some legal teams do not apply here. Verified against the Hub on 2026-09-06; if you are working from guidance written before Gemma 4, it is out of date.

What this means for your business

Google's open model family, and as of Gemma 4 it is Apache 2.0 rather than a custom Google licence.

Why it should matter to you

Licences move in both directions, and this one moved in your favour. Gemma 3 carried terms that let Google restrict use remotely and had to be passed to anyone you redistributed to, which was enough for some legal teams to refuse it outright. Gemma 4 dropped that. If your organisation rejected Gemma on licensing grounds, that decision was made against terms that no longer apply.

How it connects to our work

We re-read licences rather than carrying forward last year's answer, because the family name is not the licence. In the same week we found Gemma loosening, GLM-5.3 tightened — the flagship left MIT while its Flash sibling kept it, and Qwen's largest models left Apache 2.0 while the 27B kept it.

Parameters31B
Active per token31B (dense)
Context262K
Modalitytext, vision
Memory @ Q4_K_M~18 GB
Memory @ Q8_0~33 GB
LicenceApache 2.0
Last verified2026-09

Advantages

  • Apache 2.0 and ungated, where Gemma 3 was custom-licensed and required requesting access
  • Excellent quality per parameter among Western open models
  • Vision capable, 140+ languages
  • Free tier on OpenRouter for both the 31B and the 26B-A4B MoE variant

Disadvantages

  • Context quality degrades well before the 262K headline figure
  • Smaller fine-tune ecosystem than Qwen
  • Google has deprecated open model lines before, so plan a fallback

Reach for it when

Multilingual and vision work on a single card, now with a licence that clears OSI-style legal review. One of the strongest Western open options at this size.

Where it falls down

Long-context work at anything near the advertised 262K — test at your real document lengths. Also do not assume older Gemma guidance applies: the licence changed under the same family name.

Running it

24 GB GPU at Q4_K_M, or a 32 GB Mac. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.

will it fitGemma 4 31B against common GPUs
Gemma 4 31B at Q4_K_M
0 GB
Gemma 4 31B at Q8_0
0 GB
RTX 3060 12GB
0 GB
RTX 4060 Ti 16GB
0 GB
RTX 4090
0 GB
RTX 5090
0 GB
RTX 6000 Ada
0 GB
A100 80GB
0 GB
H100 80GB
0 GB
Weight sizes for Gemma 4 31B as recorded in this directory; GPU memory from the hardware table on the directory hub. Weights only: context adds KV cache.
where the weights fitGPU and Mac memory, checked
Q4_K_MQ8_0
RTX 3060 12GB (12 GB)
RTX 4060 Ti 16GB (16 GB)
RTX 4090 (24 GB)
RTX 5090 (32 GB)
RTX 6000 Ada (48 GB)
A100 80GB (80 GB)
H100 80GB (80 GB)
Mac, 16 GB unified
Mac, 24 GB unified
Mac, 32 GB unified
Mac, 36 GB unified~
Mac, 64 GB unified
Mac, 96 GB unified
Mac, 128 GB unified
Mac, 192 GB unified
Mac, 512 GB unified
✓ fits with room for context, ~ fits with under 15% headroom, ✕ does not fit. Weights only, computed from the figures on this page.

Free tiers carrying this model

ProviderTypeThe catchLive limits
OpenRouterAggregatorFree-pool membership changes without notice — a model you built on can stop being free, and the list above will drift. Data passes through OpenRouter and then the upstream provider, so check the privacy and routing settings before sending anything sensitive.check →
NVIDIA NIMInference hostIt is a credit grant, not a permanent free tier. When the credits run out you are on a paid plan or self-hosting the container.check →
Cloudflare Workers AIEdge inferenceThe Neuron accounting makes it hard to predict how many actual requests you get, because different models consume at very different rates.check →
Google AI Studio / Gemini APIFirst-party APIThis is the important one: on the free tier Google may use your prompts and responses to improve its products. Paid tiers do not. If you are sending anything confidential, the free tier is the wrong place for it.check →

Jurisdiction: United States

Best raw capability and the deepest tooling ecosystem. For EU personal data you are relying on a transfer framework rather than on data never leaving the bloc, so check whether your DPA and your customers accept that.

Watch for: Enterprise API tiers usually promise no training on your data; consumer tiers and free tiers frequently do not. The free tier is where this bites.

Frequently asked

Can I use Gemma 4 31B commercially?

Apache 2.0 — and this is a change worth knowing about. Gemma 3 shipped under Google's custom Gemma Terms of Use behind a manual access gate. Gemma 4 is tagged apache-2.0 on the Hub and is ungated. The remote-restriction clause and the flow-down obligation that made older Gemma unacceptable to some legal teams do not apply here. Verified against the Hub on 2026-09-06; if you are working from guidance written before Gemma 4, it is out of date.

What hardware do I need to run Gemma 4 31B?

24 GB GPU at Q4_K_M, or a 32 GB Mac. Weights alone are roughly 18 GB at Q4_K_M and 33 GB at Q8_0. Add KV cache on top of that, which grows with your context length.

What licence is Gemma 4 31B released under?

Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.

Where can I use Gemma 4 31B for free?

Free tiers carrying it include OpenRouter, NVIDIA NIM, Cloudflare Workers AI, Google AI Studio / Gemini API. Limits differ per provider and change often, so check each provider's own limits page. You can also self-host the weights, which has no rate limit at all.

Similar models

Wiring Gemma 4 31B into something real?

We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.

Maps to AI agent development, AI automation development and AI strategy and consulting. Or see it working: our case studies.