homeservicesworkaboutblogfree templatescontactFree Tools →Free AI ModelsResearch LibraryROI CalculatorSavings CalculatorAI Readiness ScoreHire vs. AutomateAutomation Quote
book a 30-min call →
home / free-models / gemma-4-26b-a4b

Gemma 4 26B A4B

Google DeepMind · United States · Apache 2.0

Commercial use: Yes — free for commercial use

Apache 2.0 and ungated, a genuine change from Gemma 3, which shipped under Google's own Gemma Terms of Use behind an access gate. Older articles still describe Gemma as restricted; for Gemma 4 that is no longer true.

Parameters25.8B total
Active per token3.8B per token
Context256K
Modalitytext, image
Memory @ Q4_K_M~17 GB
Memory @ Q8_0~27.3 GB
LicenceApache 2.0
Last verified2026-10

Unsloth's UD-Q4_K_XL is 17.0 GB. Google's quantisation-aware (QAT) 4-bit build is smaller, at 14.3 GB, because the model was trained to tolerate 4-bit weights.

Advantages

  • Cheapest capable model to call hosted: about $0.30 per million output tokens on OpenRouter in October 2026, with a free variant listed as well
  • Google reports 82.6 on MMLU-Pro and 68.2 on Tau2, with only 3.8B parameters read per token
  • A QAT 4-bit build trained for low precision, so the 4-bit file loses less than a conventional post-training quantisation
  • Covers 140+ languages, a real advantage for customer-facing work outside English

Disadvantages

  • Weak at agentic coding on the only published SWE-bench Verified figure: 17.4, from Qwen's comparison table. Google's own card does not report SWE-bench at all
  • No audio input at this size; audio is limited to the smaller E2B, E4B and 12B models
  • At 17 GB it competes for the same 24 GB card as Qwen's 27B, which is much stronger at code

Reach for it when

Multilingual chat, summarisation and image understanding at high throughput, on one 24 GB card or a 24 GB Mac.

Where it falls down

Autonomous coding agents. Use a Qwen 27B or a Qwen coder model for that, and keep Gemma for language and vision work.

Running it

24 GB GPU or 24 GB Mac at 4-bit (17.0 GB, or 14.3 GB for the QAT build). 8-bit needs a 32 GB card. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.

will it fitGemma 4 26B A4B against common GPUs
Gemma 4 26B A4B at Q4_K_M
0 GB
Gemma 4 26B A4B at Q8_0
0 GB
RTX 3060 12GB
0 GB
RTX 4060 Ti 16GB
0 GB
RTX 4090
0 GB
RTX 5090
0 GB
RTX 6000 Ada
0 GB
A100 80GB
0 GB
H100 80GB
0 GB
Weight sizes for Gemma 4 26B A4B as recorded in this directory; GPU memory from the hardware table on the directory hub. Weights only: context adds KV cache.
where the weights fitGPU and Mac memory, checked
Q4_K_MQ8_0
RTX 3060 12GB (12 GB)✕✕
RTX 4060 Ti 16GB (16 GB)✕✕
RTX 4090 (24 GB)✓✕
RTX 5090 (32 GB)✓~
RTX 6000 Ada (48 GB)✓✓
A100 80GB (80 GB)✓✓
H100 80GB (80 GB)✓✓
Mac, 16 GB unified✕✕
Mac, 24 GB unified✓✕
Mac, 32 GB unified✓~
Mac, 36 GB unified✓✓
Mac, 64 GB unified✓✓
Mac, 96 GB unified✓✓
Mac, 128 GB unified✓✓
Mac, 192 GB unified✓✓
Mac, 512 GB unified✓✓
✓ fits with room for context, ~ fits with under 15% headroom, ✕ does not fit. Weights only, computed from the figures on this page.

Which cloud instance actually runs this, and what it costs

At Q4_K_M the weights are about 17 GB, so 20 catalogued instances fit with room for context, of which the 8 best value are shown. It is a mixture-of-experts model: all 17 GB must sit in memory, but each token reads only about 2.1 GB of active experts. The speeds below assume every weight is read, so treat them as a conservative floor and the costs as a ceiling. Measured single-user speeds for this kind of model run several times higher: about 100 tok/s for gpt-oss-120b on one A100 under vLLM, and around 240 tok/s for Qwen3.6 35B-A3B with multi-token prediction on an RTX 6000 GPU in Unsloth's tests. Cheapest per token is Hetzner GEX45 (~$2.04 per million output tokens at an estimated 40 tok/s). Fastest on a single card is Azure NCads H100 v5 at roughly 197 tok/s for $9.84 per million tokens, which is the speed-versus-cost trade in one line.

InstanceGPUVRAMBandwidthEst. tok/s$/hr$/M tokens
Hetzner GEX45RTX PRO 4000 Blackwell24 GB672 GB/s~40$0.290$2.04
RunPod A100 PCIeA100 80GB80 GB1935 GB/s~114$1.190$2.90
Hetzner GEX131RTX PRO 6000 Blackwell Max-Q96 GB1792 GB/s~105$1.420$3.74
RunPod L40SL40S48 GB864 GB/s~51$0.790$4.32
RunPod H100 PCIeH100 PCIe80 GB2000 GB/s~118$1.990$4.70
Lambda A100 40GBA100 40GB40 GB1555 GB/s~91$1.990$6.04
Lambda H100 PCIeH100 PCIe80 GB2000 GB/s~118$3.290$7.77
AWS g5.xlargeA10G24 GB600 GB/s~35$1.006$7.92

Token generation on a single request is memory-bandwidth-bound, not compute-bound: the GPU must read every active weight from memory for each token it emits. So the ceiling is roughly GPU memory bandwidth divided by the size of the active weights. A 16 GB model on a 300 GB/s L4 tops out near 19 tokens/sec; the same model on a 2039 GB/s A100 tops out near 127. Figures on model pages are that arithmetic, computed from the bandwidth and weight-size columns here. They are a ceiling for one request at batch size 1, before serving overhead, so treat them as an upper bound for comparing instances rather than a benchmark. Batching raises total throughput well above this and does not raise per-request speed.

For identical silicon, the hyperscalers charge roughly 2 to 5 times what the specialist GPU clouds charge. An A100 80GB is about $3.43/GPU-hr on AWS, $3.67 on Azure and $5.03 on Google Cloud, against $1.19-1.39 on RunPod and $1.99 on Lambda. That spread is the single largest cost lever in a self-hosted inference budget, and it is larger than any saving from picking a smaller model.

Three costs that are absent from every headline rate: egress, storage for the weights, and idle time. A 70B model at Q4 is tens of gigabytes that must be pulled to the instance on every cold start, and an endpoint billed hourly is billed while it sits idle waiting for a request. Rates read 2026-09-06; check each provider's live pricing on the directory hub.

Free tiers carrying this model

ProviderTypeThe catchLive limits
OpenRouterAggregatorFree-pool membership changes without notice — a model you built on can stop being free, and the list above will drift. Data passes through OpenRouter and then the upstream provider, so check the privacy and routing settings before sending anything sensitive.check →

Jurisdiction: United States

Best raw capability and the deepest tooling ecosystem. For EU personal data you are relying on a transfer framework rather than on data never leaving the bloc, so check whether your DPA and your customers accept that.

Watch for: Enterprise API tiers usually promise no training on your data; consumer tiers and free tiers frequently do not. The free tier is where this bites.

Frequently asked

Can I use Gemma 4 26B A4B commercially?

Apache 2.0 and ungated, a genuine change from Gemma 3, which shipped under Google's own Gemma Terms of Use behind an access gate. Older articles still describe Gemma as restricted; for Gemma 4 that is no longer true.

What hardware do I need to run Gemma 4 26B A4B?

24 GB GPU or 24 GB Mac at 4-bit (17.0 GB, or 14.3 GB for the QAT build). 8-bit needs a 32 GB card. Weights alone are roughly 17 GB at Q4_K_M and 27.3 GB at Q8_0. Unsloth's UD-Q4_K_XL is 17.0 GB. Google's quantisation-aware (QAT) 4-bit build is smaller, at 14.3 GB, because the model was trained to tolerate 4-bit weights.. Add KV cache on top of that, which grows with your context length.

What licence is Gemma 4 26B A4B released under?

Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.

Where can I use Gemma 4 26B A4B for free?

Free tiers carrying it include OpenRouter. Limits differ per provider and change often, so check each provider's own limits page. You can also self-host the weights, which has no rate limit at all.

Similar models

Wiring Gemma 4 26B A4B into something real?

We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.

Maps to AI graphic design and AI automation development. Or see it working: our case studies.