Qwen3 Coder Next
Alibaba · China · Apache 2.0
Commercial use: Yes — free for commercial use
Apache 2.0, no conditions, ungated. Qwen explicitly targets integration with third-party coding tools such as Claude Code, Cline and Kilo, and the licence places no restriction on that.
Unsloth's Q4_K_M is 48.5 GB and Q8_0 is 84.8 GB. Unsloth recommends at least 3-bit; its 1-bit and 2-bit dynamic quants (19 to 27 GB) exist but cost real quality.
Advantages
- Qwen reports 74.2% on SWE-bench Verified with about 3B parameters read per token (10 of 512 experts), which is why it feels fast once loaded
- Built for coding agents: non-thinking, 256K context, and tuned for recovering from failed tool calls and test runs
- Hybrid linear attention (three Gated DeltaNet layers for every full-attention layer) keeps long-context memory growth low
- Hosted at about $0.80 per million output tokens on OpenRouter in October 2026, a quarter of the dense Qwen3.6 27B
Disadvantages
- 48.5 GB at 4-bit puts it out of reach of any single consumer GPU; it needs a large Mac, a 64 GB+ workstation card, or two 32 GB GPUs
- On the same SWE-bench Verified measure, Qwen3.6 27B reports 77.2 in 16.8 GB, so on a 24 GB budget the dense model is the better buy
- No reasoning mode, so very hard problems that benefit from deliberate thinking go to a thinking model instead
Reach for it when
A shared coding-agent endpoint on a 96 GB Mac or a 64 GB+ GPU, where its speed per token serves several developers or agents at once.
Where it falls down
Single-GPU desktops. If the model does not fit, offloading experts to system RAM erodes the speed that is its whole advantage.
Running it
A 96 GB Mac is comfortable at 4-bit (48.5 GB); a 64 GB Mac is tight once macOS keeps its share of memory. On NVIDIA, a 64 GB+ card or two 32 GB GPUs. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.
Which cloud instance actually runs this, and what it costs
At Q4_K_M the weights are about 48.5 GB, so 10 catalogued instances fit with room for context, of which the 8 best value are shown. It is a mixture-of-experts model: all 48.5 GB must sit in memory, but each token reads only about 1.7 GB of active experts. The speeds below assume every weight is read, so treat them as a conservative floor and the costs as a ceiling. Measured single-user speeds for this kind of model run several times higher: about 100 tok/s for gpt-oss-120b on one A100 under vLLM, and around 240 tok/s for Qwen3.6 35B-A3B with multi-token prediction on an RTX 6000 GPU in Unsloth's tests. Cheapest per token is RunPod A100 PCIe (~$8.29 per million output tokens at an estimated 40 tok/s). Fastest on a single card is Azure NCads H100 v5 at roughly 69 tok/s for $28.07 per million tokens, which is the speed-versus-cost trade in one line.
| Instance | GPU | VRAM | Bandwidth | Est. tok/s | $/hr | $/M tokens |
|---|---|---|---|---|---|---|
| RunPod A100 PCIe | A100 80GB | 80 GB | 1935 GB/s | ~40 | $1.190 | $8.29 |
| Hetzner GEX131 | RTX PRO 6000 Blackwell Max-Q | 96 GB | 1792 GB/s | ~37 | $1.420 | $10.68 |
| RunPod H100 PCIe | H100 PCIe | 80 GB | 2000 GB/s | ~41 | $1.990 | $13.40 |
| Lambda H100 PCIe | H100 PCIe | 80 GB | 2000 GB/s | ~41 | $3.290 | $22.16 |
| Azure NC24ads A100 v4 | A100 80GB | 80 GB | 2039 GB/s | ~42 | $3.673 | $24.27 |
| Azure NCads H100 v5 | H100 | 94 GB | 3350 GB/s | ~69 | $6.980 | $28.07 |
| AWS g6e.12xlarge | 4x L40S | 192 GB | 864 GB/s | ~18 | $10.490 | $163.57 |
| AWS p5.48xlarge | 8x H100 | 640 GB | 3350 GB/s | ~69 | $55.040 | $221.35 |
Token generation on a single request is memory-bandwidth-bound, not compute-bound: the GPU must read every active weight from memory for each token it emits. So the ceiling is roughly GPU memory bandwidth divided by the size of the active weights. A 16 GB model on a 300 GB/s L4 tops out near 19 tokens/sec; the same model on a 2039 GB/s A100 tops out near 127. Figures on model pages are that arithmetic, computed from the bandwidth and weight-size columns here. They are a ceiling for one request at batch size 1, before serving overhead, so treat them as an upper bound for comparing instances rather than a benchmark. Batching raises total throughput well above this and does not raise per-request speed.
For identical silicon, the hyperscalers charge roughly 2 to 5 times what the specialist GPU clouds charge. An A100 80GB is about $3.43/GPU-hr on AWS, $3.67 on Azure and $5.03 on Google Cloud, against $1.19-1.39 on RunPod and $1.99 on Lambda. That spread is the single largest cost lever in a self-hosted inference budget, and it is larger than any saving from picking a smaller model.
Three costs that are absent from every headline rate: egress, storage for the weights, and idle time. A 70B model at Q4 is tens of gigabytes that must be pulled to the instance on every cold start, and an endpoint billed hourly is billed while it sits idle waiting for a request. Rates read 2026-09-06; check each provider's live pricing on the directory hub.
Jurisdiction: China
This is the crucial split most write-ups miss: using a Chinese lab's HOSTED API sends your data to Chinese infrastructure and is a real procurement question. Downloading their OPEN WEIGHTS and running them on your own hardware, or on a Western host, sends nothing anywhere. DeepSeek and Z.ai publish under MIT; Alibaba publishes most of Qwen3 under Apache 2.0. Those are among the most permissive licences on this page.
Watch for: Several US states and a number of government bodies restrict Chinese-hosted AI services on official devices. That restriction is about the hosted service, not about the weights running in your own VPC — but expect to have to explain the difference to a security reviewer.
Frequently asked
Can I use Qwen3 Coder Next commercially?
Apache 2.0, no conditions, ungated. Qwen explicitly targets integration with third-party coding tools such as Claude Code, Cline and Kilo, and the licence places no restriction on that.
What hardware do I need to run Qwen3 Coder Next?
A 96 GB Mac is comfortable at 4-bit (48.5 GB); a 64 GB Mac is tight once macOS keeps its share of memory. On NVIDIA, a 64 GB+ card or two 32 GB GPUs. Weights alone are roughly 48.5 GB at Q4_K_M and 84.8 GB at Q8_0. Unsloth's Q4_K_M is 48.5 GB and Q8_0 is 84.8 GB. Unsloth recommends at least 3-bit; its 1-bit and 2-bit dynamic quants (19 to 27 GB) exist but cost real quality.. Add KV cache on top of that, which grows with your context length.
What licence is Qwen3 Coder Next released under?
Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.
Where can I use Qwen3 Coder Next for free?
Self-hosting the weights is the free route. A 96 GB Mac is comfortable at 4-bit (48.5 GB); a 64 GB Mac is tight once macOS keeps its share of memory. On NVIDIA, a 64 GB+ card or two 32 GB GPUs.
Similar models
DeepSeek-V4-Flash
Hosted, or 256 GB+ unified memory at Q4. Not a laptop model.
GLM-5.3-Flash
80-96 GB at Q4_K_M. A workstation or Ultra-class Mac.
Qwen3.8 27B
24 GB GPU at Q4_K_M, or a 32 GB Mac at Q6_K.
Qwen3 Coder 30B A3B
24 GB GPU at Q4_K_M (18.6 GB), or a 32 GB Mac. Its 3.3B active parameters make it fast even on modest memory bandwidth.
Wiring Qwen3 Coder Next into something real?
We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.
Maps to AI agent development, AI automation development and AI strategy and consulting. Or see it working: our case studies.