Gemma 4 E4B
Google DeepMind · United States · Apache 2.0
Commercial use: Yes — free for commercial use
Apache 2.0 and ungated, like the rest of Gemma 4. That matters most at this size, because an on-device model is the one most likely to end up inside a product you distribute.
Unsloth's Q4_K_M is 5.0 GB and the QAT 4-bit build is 4.2 GB. The 'E' means effective: per-layer embeddings bring the stored size to about 8B parameters while compute behaves like a 4B model.
Advantages
- Takes audio as a direct input alongside text and images, so speech can go straight into a local model without a separate transcription step
- Small enough for a laptop: the E2B and E4B models are the Gemma 4 sizes designed for phones and laptops
- Apache 2.0 and ungated, so it can be bundled into an app without Google's former Gemma terms
- A 4.2 GB QAT build trained to hold its quality at 4-bit
Disadvantages
- Weak at agentic work: Google reports 42.2 on Tau2 and 52.0 on LiveCodeBench v6, far below the 26B MoE
- Context tops out at 128K, half the larger Gemma 4 and Qwen models
- The 'E4B' name undersells its memory footprint. Plan for 8B-class storage, not 4B
Reach for it when
On-device voice notes, meeting snippets and image questions, where the audio or photo should never leave the machine.
Where it falls down
Tool-calling agents, code, and long-document analysis. Pair it with a larger model rather than asking it to do everything.
Running it
Any 8 GB GPU or 16 GB laptop at Q4_K_M. Unsloth sizes 4-bit at about 5.5 to 6 GB of total memory. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.
Which cloud instance actually runs this, and what it costs
At Q4_K_M the weights are about 5 GB, so 20 catalogued instances fit with room for context, of which the 8 best value are shown. Cheapest per token is Hetzner GEX45 (~$0.60 per million output tokens at an estimated 134 tok/s). Fastest on a single card is Azure NCads H100 v5 at roughly 670 tok/s for $2.89 per million tokens, which is the speed-versus-cost trade in one line.
| Instance | GPU | VRAM | Bandwidth | Est. tok/s | $/hr | $/M tokens |
|---|---|---|---|---|---|---|
| Hetzner GEX45 | RTX PRO 4000 Blackwell | 24 GB | 672 GB/s | ~134 | $0.290 | $0.60 |
| RunPod A100 PCIe | A100 80GB | 80 GB | 1935 GB/s | ~387 | $1.190 | $0.85 |
| Hetzner GEX131 | RTX PRO 6000 Blackwell Max-Q | 96 GB | 1792 GB/s | ~358 | $1.420 | $1.10 |
| RunPod L40S | L40S | 48 GB | 864 GB/s | ~173 | $0.790 | $1.27 |
| RunPod H100 PCIe | H100 PCIe | 80 GB | 2000 GB/s | ~400 | $1.990 | $1.38 |
| Lambda A100 40GB | A100 40GB | 40 GB | 1555 GB/s | ~311 | $1.990 | $1.78 |
| Lambda H100 PCIe | H100 PCIe | 80 GB | 2000 GB/s | ~400 | $3.290 | $2.28 |
| AWS g5.xlarge | A10G | 24 GB | 600 GB/s | ~120 | $1.006 | $2.33 |
Token generation on a single request is memory-bandwidth-bound, not compute-bound: the GPU must read every active weight from memory for each token it emits. So the ceiling is roughly GPU memory bandwidth divided by the size of the active weights. A 16 GB model on a 300 GB/s L4 tops out near 19 tokens/sec; the same model on a 2039 GB/s A100 tops out near 127. Figures on model pages are that arithmetic, computed from the bandwidth and weight-size columns here. They are a ceiling for one request at batch size 1, before serving overhead, so treat them as an upper bound for comparing instances rather than a benchmark. Batching raises total throughput well above this and does not raise per-request speed.
For identical silicon, the hyperscalers charge roughly 2 to 5 times what the specialist GPU clouds charge. An A100 80GB is about $3.43/GPU-hr on AWS, $3.67 on Azure and $5.03 on Google Cloud, against $1.19-1.39 on RunPod and $1.99 on Lambda. That spread is the single largest cost lever in a self-hosted inference budget, and it is larger than any saving from picking a smaller model.
Three costs that are absent from every headline rate: egress, storage for the weights, and idle time. A 70B model at Q4 is tens of gigabytes that must be pulled to the instance on every cold start, and an endpoint billed hourly is billed while it sits idle waiting for a request. Rates read 2026-09-06; check each provider's live pricing on the directory hub.
Jurisdiction: United States
Best raw capability and the deepest tooling ecosystem. For EU personal data you are relying on a transfer framework rather than on data never leaving the bloc, so check whether your DPA and your customers accept that.
Watch for: Enterprise API tiers usually promise no training on your data; consumer tiers and free tiers frequently do not. The free tier is where this bites.
Frequently asked
Can I use Gemma 4 E4B commercially?
Apache 2.0 and ungated, like the rest of Gemma 4. That matters most at this size, because an on-device model is the one most likely to end up inside a product you distribute.
What hardware do I need to run Gemma 4 E4B?
Any 8 GB GPU or 16 GB laptop at Q4_K_M. Unsloth sizes 4-bit at about 5.5 to 6 GB of total memory. Weights alone are roughly 5 GB at Q4_K_M and 8.3 GB at Q8_0. Unsloth's Q4_K_M is 5.0 GB and the QAT 4-bit build is 4.2 GB. The 'E' means effective: per-layer embeddings bring the stored size to about 8B parameters while compute behaves like a 4B model.. Add KV cache on top of that, which grows with your context length.
What licence is Gemma 4 E4B released under?
Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.
Where can I use Gemma 4 E4B for free?
Self-hosting the weights is the free route. Any 8 GB GPU or 16 GB laptop at Q4_K_M. Unsloth sizes 4-bit at about 5.5 to 6 GB of total memory.
Similar models
Gemma 4 26B A4B
24 GB GPU or 24 GB Mac at 4-bit (17.0 GB, or 14.3 GB for the QAT build). 8-bit needs a 32 GB card.
Qwen3.6 27B
24 GB GPU at Q4_K_M, or a 24 GB Mac with modest context. Add about 2 GB for the MTP build.
Qwen3.6 35B A3B
32 GB GPU or 32 GB Mac at 4-bit. On 24 GB, offload some experts to system RAM and accept slower generation.
Qwen3.5 9B
8 GB GPU at Q4_K_M, a 12 GB GPU at Q8_0, or any 16 GB Mac.
Wiring Gemma 4 E4B into something real?
We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.
Maps to voice AI development and AI automation development. Or see it working: our case studies.