homeservicesworkaboutblogfree templatescontactFree Tools →Free AI ModelsResearch LibraryROI CalculatorSavings CalculatorAI Readiness ScoreHire vs. AutomateAutomation Quote
book a 30-min call →
home / free-models / qwen3-6-27b

Qwen3.6 27B

Alibaba · China · Apache 2.0

Commercial use: Yes — free for commercial use

Apache 2.0, no conditions, and ungated on the Hub. The same licence carries through to Unsloth's GGUF, MTP and NVFP4 builds, so whichever format you pick, nothing about your legal position changes.

Parameters27.8B
Active per token27.8B (dense)
Context256K
Modalitytext, image, code
Memory @ Q4_K_M~16.8 GB
Memory @ Q8_0~28.6 GB
LicenceApache 2.0
Last verified2026-10

Unsloth's Q4_K_M is 16.8 GB. The MTP build adds roughly 2 GB, and Unsloth's own sizing puts the 4-bit MTP variant at about 19 GB of total memory.

Advantages

  • Qwen reports 77.2 on SWE-bench Verified, ahead of its own 397B-A17B flagship from the previous generation (76.2) on the same table
  • Three of every four layers use linear attention, so the memory needed for long context grows far more slowly than in a conventional 27B model
  • Ships with a built-in multi-token prediction head, which Unsloth's MTP builds use for roughly 1.4x to 2.2x faster generation without changing the output
  • Reads images as well as text, so screenshots, scanned forms and diagrams are inputs rather than a separate pipeline

Disadvantages

  • Superseded in August 2026 by Qwen3.8 27B, which has exactly the same architecture and parameter count and scores higher on Qwen's own comparison (SWE-bench Pro 61.7 against 53.5)
  • Dense, so every token reads all 27.8B parameters. That is why hosted providers charge about $3.25 per million output tokens for it, triple the MoE sibling
  • Thinking mode is on by default and roughly multiplies output length on simple requests

Reach for it when

Teams that already validated prompts and tool calls against Qwen3.6 and want to keep a known-good model, or anyone with a Blackwell GPU using Unsloth's NVFP4 build.

Where it falls down

New deployments with no history on 3.6. Qwen3.8 27B is a drop-in replacement with the same memory footprint and better scores, so starting on 3.6 today gains nothing.

Running it

24 GB GPU at Q4_K_M, or a 24 GB Mac with modest context. Add about 2 GB for the MTP build. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.

will it fitQwen3.6 27B against common GPUs
Qwen3.6 27B at Q4_K_M
0 GB
Qwen3.6 27B at Q8_0
0 GB
RTX 3060 12GB
0 GB
RTX 4060 Ti 16GB
0 GB
RTX 4090
0 GB
RTX 5090
0 GB
RTX 6000 Ada
0 GB
A100 80GB
0 GB
H100 80GB
0 GB
Weight sizes for Qwen3.6 27B as recorded in this directory; GPU memory from the hardware table on the directory hub. Weights only: context adds KV cache.
where the weights fitGPU and Mac memory, checked
Q4_K_MQ8_0
RTX 3060 12GB (12 GB)✕✕
RTX 4060 Ti 16GB (16 GB)✕✕
RTX 4090 (24 GB)✓✕
RTX 5090 (32 GB)✓~
RTX 6000 Ada (48 GB)✓✓
A100 80GB (80 GB)✓✓
H100 80GB (80 GB)✓✓
Mac, 16 GB unified✕✕
Mac, 24 GB unified✓✕
Mac, 32 GB unified✓~
Mac, 36 GB unified✓✓
Mac, 64 GB unified✓✓
Mac, 96 GB unified✓✓
Mac, 128 GB unified✓✓
Mac, 192 GB unified✓✓
Mac, 512 GB unified✓✓
✓ fits with room for context, ~ fits with under 15% headroom, ✕ does not fit. Weights only, computed from the figures on this page.

Which cloud instance actually runs this, and what it costs

At Q4_K_M the weights are about 16.8 GB, so 20 catalogued instances fit with room for context, of which the 8 best value are shown. Cheapest per token is Hetzner GEX45 (~$2.01 per million output tokens at an estimated 40 tok/s). Fastest on a single card is Azure NCads H100 v5 at roughly 199 tok/s for $9.72 per million tokens, which is the speed-versus-cost trade in one line.

InstanceGPUVRAMBandwidthEst. tok/s$/hr$/M tokens
Hetzner GEX45RTX PRO 4000 Blackwell24 GB672 GB/s~40$0.290$2.01
RunPod A100 PCIeA100 80GB80 GB1935 GB/s~115$1.190$2.87
Hetzner GEX131RTX PRO 6000 Blackwell Max-Q96 GB1792 GB/s~107$1.420$3.70
RunPod L40SL40S48 GB864 GB/s~51$0.790$4.27
RunPod H100 PCIeH100 PCIe80 GB2000 GB/s~119$1.990$4.64
Lambda A100 40GBA100 40GB40 GB1555 GB/s~93$1.990$5.97
Lambda H100 PCIeH100 PCIe80 GB2000 GB/s~119$3.290$7.68
AWS g5.xlargeA10G24 GB600 GB/s~36$1.006$7.82

Token generation on a single request is memory-bandwidth-bound, not compute-bound: the GPU must read every active weight from memory for each token it emits. So the ceiling is roughly GPU memory bandwidth divided by the size of the active weights. A 16 GB model on a 300 GB/s L4 tops out near 19 tokens/sec; the same model on a 2039 GB/s A100 tops out near 127. Figures on model pages are that arithmetic, computed from the bandwidth and weight-size columns here. They are a ceiling for one request at batch size 1, before serving overhead, so treat them as an upper bound for comparing instances rather than a benchmark. Batching raises total throughput well above this and does not raise per-request speed.

For identical silicon, the hyperscalers charge roughly 2 to 5 times what the specialist GPU clouds charge. An A100 80GB is about $3.43/GPU-hr on AWS, $3.67 on Azure and $5.03 on Google Cloud, against $1.19-1.39 on RunPod and $1.99 on Lambda. That spread is the single largest cost lever in a self-hosted inference budget, and it is larger than any saving from picking a smaller model.

Three costs that are absent from every headline rate: egress, storage for the weights, and idle time. A 70B model at Q4 is tens of gigabytes that must be pulled to the instance on every cold start, and an endpoint billed hourly is billed while it sits idle waiting for a request. Rates read 2026-09-06; check each provider's live pricing on the directory hub.

Jurisdiction: China

This is the crucial split most write-ups miss: using a Chinese lab's HOSTED API sends your data to Chinese infrastructure and is a real procurement question. Downloading their OPEN WEIGHTS and running them on your own hardware, or on a Western host, sends nothing anywhere. DeepSeek and Z.ai publish under MIT; Alibaba publishes most of Qwen3 under Apache 2.0. Those are among the most permissive licences on this page.

Watch for: Several US states and a number of government bodies restrict Chinese-hosted AI services on official devices. That restriction is about the hosted service, not about the weights running in your own VPC — but expect to have to explain the difference to a security reviewer.

The paper behind it

This model has a published technical report: arXiv:2505.09388 It is indexed in our research library alongside the work it builds on.

Frequently asked

Can I use Qwen3.6 27B commercially?

Apache 2.0, no conditions, and ungated on the Hub. The same licence carries through to Unsloth's GGUF, MTP and NVFP4 builds, so whichever format you pick, nothing about your legal position changes.

What hardware do I need to run Qwen3.6 27B?

24 GB GPU at Q4_K_M, or a 24 GB Mac with modest context. Add about 2 GB for the MTP build. Weights alone are roughly 16.8 GB at Q4_K_M and 28.6 GB at Q8_0. Unsloth's Q4_K_M is 16.8 GB. The MTP build adds roughly 2 GB, and Unsloth's own sizing puts the 4-bit MTP variant at about 19 GB of total memory.. Add KV cache on top of that, which grows with your context length.

What licence is Qwen3.6 27B released under?

Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.

Where can I use Qwen3.6 27B for free?

Self-hosting the weights is the free route. 24 GB GPU at Q4_K_M, or a 24 GB Mac with modest context. Add about 2 GB for the MTP build.

Similar models

Wiring Qwen3.6 27B into something real?

We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.

Maps to AI graphic design and AI automation development. Or see it working: our case studies.