homeservicesworkaboutblogfree templatescontactFree Tools →Free AI ModelsResearch LibraryROI CalculatorSavings CalculatorAI Readiness ScoreHire vs. AutomateAutomation Quote
book a 30-min call →
home / blog / The Most Downloaded Unsloth Models: Hardware, Cloud Costs and Which to Run in 2026

The Most Downloaded Unsloth Models: Hardware, Cloud Costs and Which to Run in 2026

We re-counted the most downloaded Unsloth models against the Hugging Face API. The viral chart misses the real number one, eleven models hide behind fifteen rows, and renting a GPU for one user usually costs more per token than the API. Here is what each model needs, costs and is good for.

The Most Downloaded Unsloth Models: Hardware, Cloud Costs and Which to Run in 2026

A chart titled "Most downloaded Unsloth models of all time" has been doing the rounds, and it is a useful starting point for anyone choosing an open model to run themselves. We checked it against the Hugging Face API on 5 October 2026 before writing a word. The ranking is close but stale, it leaves out the single most downloaded Unsloth repository by a wide margin, and its fifteen rows hide only eleven distinct models. More usefully for a business, the chart says nothing about what each one needs to run, what it costs, or what it is good for. This guide covers exactly that, model by model, with the hardware, the cloud bill, and a recommendation at the end.

Metric Verified figure (5 October 2026)
Most downloaded Unsloth repository Qwen3-Coder-30B-A3B GGUF, 28.6M all-time
Distinct models in the top 15 11 (several rows are the same model in another format)
Cheapest capable hosted model in the set Gemma 4 26B A4B, $0.30 per million output tokens
Cheapest single-user self-hosting of Qwen3.6 27B About $2.01 per million tokens on a Hetzner GEX45

What the Most Downloaded Unsloth Models Actually Are

Unsloth is an open-source project best known for making fine-tuning faster and lighter, and for publishing quantised versions of new open models within hours of release. When a business owner downloads "Qwen3.6 27B" through LM Studio or Ollama's Hugging Face integration, the file very often comes from Unsloth's account rather than from Alibaba's. That is why Unsloth's download counts work as a rough census of which open models people are really trying to run locally.

the recountWhat the download chart actually says, once you check it
0.0Mall-time downloads of the real number one, Qwen3 Coder 30B A3B, missing from the viral chart
0distinct models behind the fifteen rows of the top 15
0 of 15top-15 repositories under Apache 2.0
0.0%of Llama 3.1 8B 4-bit's all-time downloads that happened in the last 30 days
Hugging Face API, all 1,460 public repositories under huggingface.co/unsloth, downloadsAllTime and 30-day downloads read on 5 October 2026.

The headline finding from the recount is simple. The viral chart lists Qwen3.8 27B as number one at 14.5M downloads. The Hugging Face API shows Qwen3.8 27B's GGUF build has since climbed to 16.9M, and that Qwen3-Coder-30B-A3B-Instruct-GGUF sits above it at 28.6M, with 7.4M of those in the last 30 days alone. Whatever filter produced the chart, it dropped the model that more people download than any other Unsloth upload. The rest of the order is close, with neighbouring positions swapped by a few hundred thousand downloads as the counts have kept moving.

The Verified Top 15 Unsloth Models

This is the ranking rebuilt from all 1,460 public Unsloth repositories, sorted by Hugging Face's downloadsAllTime field on 5 October 2026.

verified rankingAll-time downloads, top 15 Unsloth repositories
Qwen3-Coder-30B-A3B GGUF
0
Qwen3.8-27B GGUF
0
gemma-4-26B-A4B GGUF
0
Qwen3.6-27B NVFP4
0
Llama-3.1-8B bnb-4bit
0
Qwen3.5-9B GGUF
0
Qwen3.6-35B-A3B GGUF
0
DeepSeek-R1 GGUF
0
Qwen3.6-27B MTP GGUF
0
Llama-3.1-8B (16-bit)
0
Qwen3.5-4B GGUF
0
Mistral-7B v0.3 bnb-4bit
0
Qwen3.6-27B GGUF
0
gemma-4-E4B GGUF
0
Qwen3.8-27B NVFP4
0
Hugging Face API downloadsAllTime, read 5 October 2026. A download is any GET or HEAD request to a counted file, so these are fetches, not users.
Rank Repository All-time Last 30 days Base model Format
1 Qwen3-Coder-30B-A3B-Instruct-GGUF 28.57M 7.36M Qwen3 Coder 30B A3B GGUF
2 Qwen3.8-27B-GGUF 16.87M 6.55M Qwen3.8 27B GGUF
3 gemma-4-26B-A4B-it-GGUF 11.17M 0.53M Gemma 4 26B A4B GGUF
4 Qwen3.6-27B-NVFP4 9.72M 0.96M Qwen3.6 27B NVFP4
5 Meta-Llama-3.1-8B-Instruct-bnb-4bit 9.65M 0.07M Llama 3.1 8B bnb-4bit
6 Qwen3.5-9B-GGUF 9.13M 1.24M Qwen3.5 9B GGUF
7 Qwen3.6-35B-A3B-GGUF 8.83M 1.30M Qwen3.6 35B A3B GGUF
8 DeepSeek-R1-GGUF 8.44M 0.02M DeepSeek-R1 GGUF
9 Qwen3.6-27B-MTP-GGUF 6.95M 0.75M Qwen3.6 27B MTP GGUF
10 Meta-Llama-3.1-8B-Instruct 6.85M 0.21M Llama 3.1 8B 16-bit
11 Qwen3.5-4B-GGUF 6.75M 1.06M Qwen3.5 4B GGUF
12 mistral-7b-v0.3-bnb-4bit 6.53M 0.37M Mistral 7B v0.3 bnb-4bit
13 Qwen3.6-27B-GGUF 6.16M 0.75M Qwen3.6 27B GGUF
14 gemma-4-E4B-it-GGUF 5.71M 0.62M Gemma 4 E4B GGUF
15 Qwen3.8-27B-NVFP4 5.42M 2.31M Qwen3.8 27B NVFP4

Just outside the fifteen sit Qwen3.6-35B-A3B NVFP4 (5.37M), Qwen3.5-35B-A3B (5.34M), Qwen3-Coder-Next (5.10M), Gemma 4 31B (5.06M) and gpt-oss-20b (4.94M). Qwen3-Coder-Next appears in the viral chart, so we cover it below as well.

Why Download Counts Are Not a Measure of Users

Before reading a ranking like this as "what businesses run", it is worth knowing how the number is produced. Hugging Face's own documentation on model download stats says the count is taken server-side and that every HTTP request to the counted files, including GET and HEAD, is counted as a download. For GGUF repositories every GGUF file counts, and the documentation notes this double counts when someone clones a whole repository. Excluding CI pipelines or counting unique downloaders needs the publisher's own analytics, which are not public.

Two consequences follow, and both change how you should read the chart:

  1. Automated pulls inflate the totals. A fine-tuning notebook that loads a pre-quantised checkpoint such as unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit registers a download every time it runs. That repository's own model card points readers to a free Colab notebook for fine-tuning Llama 3.1 8B, which goes a long way to explaining why a 2024 model still sits in the top five.
  2. All-time totals are history, not current demand. The Llama 3.1 8B 4-bit build has 9.65M all-time downloads but only about 68,000 in the last 30 days. DeepSeek-R1's GGUF has 8.44M all-time and about 20,000 in the last month. Both were genuinely huge, and both have been overtaken.

The better signal is the 30-day column. By that measure, the models people are pulling right now are Qwen3-Coder 30B A3B, Qwen3.8 27B, the Qwen3.6 and Qwen3.5 families, and Gemma 4.

Fifteen Rows, Eleven Models: GGUF, NVFP4, MTP and bnb-4bit Explained

Several rows in the chart are the same model in a different file. Qwen3.6 27B appears three times (NVFP4, MTP GGUF and plain GGUF) and Qwen3.8 27B twice. The format is not a detail: it decides what hardware the file will run on at all.

which file runs whereThe five formats in the chart, and what each one needs
llama.cpp / OllamavLLM / SGLangApple MacOlder NVIDIA (pre-Blackwell)Fine-tuning
GGUF (incl. Unsloth Dynamic)✓~✓✓✕
QAT GGUF (Gemma 4)✓~✓✓✕
MTP GGUF (Qwen3.6)✓✕✓✓✕
NVFP4✕✓✕✕✕
bnb-4bit✕~✕✓✓
Unsloth documentation (NVFP4, MTP, Dynamic GGUF and Qwen3.6 guides) and Hugging Face transformers bitsandbytes docs, read 5 October 2026. Part means supported with caveats or an extra step.

GGUF (and Unsloth Dynamic GGUF). The file format of llama.cpp, and so of Ollama, LM Studio and most desktop tools. It runs on CPUs, Apple Silicon, NVIDIA and AMD, and it can split a model between GPU and system RAM when it does not fit. Unsloth's "UD" files (for example UD-Q4_K_XL) are its Dynamic quants: instead of compressing every layer to 4 bits, Unsloth measures which layers are sensitive and keeps those at higher precision, so the same file size loses less quality than a uniform Q4_K_M. If you are unsure which file to pick, a UD-Q4_K_XL GGUF is the sensible default.

QAT GGUF (Gemma 4). Quantisation-aware training means Google trained the model to tolerate 4-bit weights, rather than compressing it afterwards. The practical gain is size at equal quality: Gemma 4 26B A4B's QAT 4-bit file is 14.3 GB against 17.0 GB for the standard dynamic build.

MTP GGUF. Multi-token prediction. Qwen trained a small prediction head into Qwen3.5 and Qwen3.6 that guesses the next few tokens, and the main model checks them in one pass. Because only verified tokens are kept, the output does not change. Unsloth reports roughly 1.4x to 2.2x faster generation in llama.cpp for about 2 GB of extra memory, with smaller gains on low-bandwidth machines such as older Macs.

NVFP4. NVIDIA's native 4-bit floating-point format, which runs directly on the FP4 tensor cores of Blackwell GPUs (RTX 50-series, RTX PRO 6000, DGX Spark, B200). It holds accuracy well: on Unsloth's own tests, its Qwen3.6 27B NVFP4 build scored 86.25 on MMLU-Pro against 85.96 for the full BF16 model. It needs a Blackwell GPU and runs in vLLM or SGLang, not in llama.cpp or Ollama.

bnb-4bit. The bitsandbytes 4-bit format used by Hugging Face transformers. It exists mainly for QLoRA fine-tuning on a single NVIDIA GPU. For serving, it is slower than the alternatives, which is the other reason the two 2024-era bnb-4bit models in the chart say more about fine-tuning tutorials than about production deployments.

The NVFP4 speed claim, read carefully

Unsloth headlines its Dynamic NVFP4 Qwen3.6 builds as running up to 2.5x faster. Its own documentation states that the benchmarks were run on one B200 serving 128 concurrent requests. On the same page, for single-person decode speed, its 27B build is listed as 1.03x faster than other NVFP4 quants, with 1.17x and 1.22x for the 35B variants. That is not a criticism of the work, which is careful and openly reported. It is a reminder that format speed-ups are usually server throughput figures. If one person is chatting with one model on one card, the gain from NVFP4 is small, and MTP will do more for you.

picking a formatWhich Unsloth file to download

If you are fine-tuning, use a bnb-4bit or 16-bit checkpoint. If you are serving many users on a Blackwell GPU with vLLM or SGLang, use NVFP4. Everyone else uses GGUF: the MTP build for Qwen3.6 if you have about 2 GB spare, the QAT build for Gemma 4 at 4-bit, and a UD-Q4_K_XL dynamic quant otherwise.

The format descriptions in this section, from Unsloth's documentation.

The Models, One by One

Each entry links to our free-models directory, where the licence verdict, hardware sizing and a ranked table of cloud instances are kept current.

Qwen3 Coder 30B A3B: the real number one

Released in July 2025, Qwen3 Coder 30B A3B is a mixture-of-experts coding model with 30.5B parameters in total and only 3.3B used per token. That combination explains its popularity: the 4-bit GGUF is 18.6 GB, so it fits a single 24 GB card, and because so little of it is read per token it answers quickly inside an editor. It is non-thinking by design, so it does not spend tokens reasoning before it writes code. Its weakness is age. Its SWE-bench Verified score is reported anywhere between roughly 49% and 60% depending on the test harness, and Qwen's newer models clear it comfortably. It remains a fast, cheap autocomplete and editing model; for autonomous agent work, look further down this list.

Qwen3.8 27B and Qwen3.6 27B: one model, two vintages

These two deserve to be read together, because the Hugging Face API shows they are the same architecture with an identical parameter count: 27,781,427,952 parameters each. Qwen3.8 27B, released in August 2026, looks like a further-trained Qwen3.6 27B rather than a new design. Both are dense, both read images as well as text, and both use linear attention in 48 of their 64 layers, which keeps long-context memory growth low.

On Qwen's own comparison table, Qwen3.8 27B scores 61.7 on SWE-bench Pro against 53.5 for Qwen3.6 27B, and 90.3 against 83.9 on LiveCodeBench v6. Qwen3.6 27B is still a very strong model in its own right: Qwen reports 77.2 on SWE-bench Verified, above the 76.2 of its own previous-generation 397B flagship. The practical rule is simple. If you are starting fresh, use Qwen3.8 27B. Stay on 3.6 only if you have prompts and tool calls validated against it and no time to re-test. The 4-bit GGUF is about 16.5 to 16.8 GB for either, which fits a 24 GB GPU, or a 24 GB Mac with modest context.

Qwen3.6 35B A3B: faster, but hungrier than it looks

Qwen3.6 35B A3B routes each token through 8 of its 256 experts, so roughly 3B parameters are read per token. Qwen reports 73.4 on SWE-bench Verified, close to the dense 27B's 77.2. In Unsloth's tests on an RTX 6000 GPU with MTP enabled, it generated about 240 tokens per second against 160 for the dense 27B.

The catch is memory. Every expert must sit in memory even though few are used, so the 4-bit file is 22.4 GB, more than the dense model. On the most common 24 GB card that leaves almost nothing for context. Plan on a 32 GB GPU such as the RTX 5090, or a 32 GB Mac, or accept some experts offloaded to system RAM.

Gemma 4 26B A4B: the cheap all-rounder that is not a coder

Gemma 4 is the first Gemma generation released under Apache 2.0 and without an access gate. Gemma 3 shipped under Google's own Gemma Terms of Use, and plenty of articles still describe Gemma that way. Gemma 4 26B A4B reads 3.8B of its 25.8B parameters per token, takes text and images, covers 140+ languages, and is the cheapest capable model here to call hosted, at about $0.30 per million output tokens on OpenRouter with a free variant listed as well.

Google reports 82.6 on MMLU-Pro and 68.2 on Tau2. Google's model card does not report SWE-bench at all. The only SWE-bench Verified figure we could find comes from Qwen's comparison table, which puts Gemma 4 26B A4B at 17.4 against 73.4 for Qwen3.6 35B A3B. That number comes from a competitor's evaluation, so weigh it accordingly. Where Qwen reproduced Google's own published numbers in the same table, they match exactly, which suggests the comparison was run in good faith. The safe reading: use Gemma 4 for multilingual chat, summarisation and image work, and keep coding agents on a Qwen model.

Qwen3.5 9B and Qwen3.5 4B: the small models doing real work

Qwen3.5 9B is 5.7 GB at 4-bit and fits an 8 GB GPU. On Qwen's table it scores 79.1 on TAU2-Bench and 66.1 on BFCL-V4, tool-use benchmarks where it beats Qwen's own 80B Qwen3-Next. Coding is its weak spot: 65.6 on LiveCodeBench v6 against 74.6 for gpt-oss-20b. Qwen3.5 4B is 2.7 GB at 4-bit, runs usefully on a CPU, and posts 79.9 on TAU2-Bench, level with the 9B. Its general reasoning drops clearly, with GPQA Diamond at 76.2 against 81.7.

The lesson is about the job, not the model. Narrow, structured work (routing a request, extracting fields, choosing a tool) does not need a 27B model. Both small models have reasoning switched off by default, so they answer quickly unless asked to think. Treat the tool-use scores as a reason to test them on your own tool schema, not as a substitute for that test.

Gemma 4 E4B: the one that hears

Gemma 4 E4B is the only model in this list that takes audio as a direct input alongside text and images. The "E" means effective: per-layer embeddings bring the stored size to about 8B parameters while it computes like a 4B model, so the 4-bit file is 5.0 GB (4.2 GB for the QAT build). It is weak at agentic work, with Google reporting 42.2 on Tau2, and its context stops at 128K. Its value is privacy at the edge. A voice note, a meeting snippet or a photo can be understood on the laptop that recorded it, without ever being uploaded.

Qwen3 Coder Next: strong, fast, and too big for one consumer card

Qwen3 Coder Next is an 80B mixture-of-experts model that uses 10 of its 512 experts per token, about 3B active parameters. Qwen reports 74.2% on SWE-bench Verified, and it is built for coding agents: non-thinking, 256K context, and tuned to recover from failed tool calls. The 4-bit GGUF is 48.5 GB, which puts it beyond any single consumer GPU. It needs a large Mac, a 64 GB+ workstation card, or two 32 GB GPUs. On a 24 GB budget, the dense Qwen3.6 or Qwen3.8 27B scores in the same range on the same benchmark in a third of the memory, so it is usually the better buy.

The 2024-era three: Llama 3.1 8B, Mistral 7B v0.3 and DeepSeek-R1

These still rank on all-time totals, but the 30-day column shows how far demand has moved on. Llama 3.1 8B is gated behind a manual approval on Meta's repository and ships under the Llama 3.1 Community License, which allows commercial use below 700 million monthly active users but is not an open-source licence. Mistral 7B v0.3 is Apache 2.0 and still a perfectly serviceable small model, but Qwen3.5 9B now does more in a similar footprint. DeepSeek-R1 is MIT licensed and historically important, but it is a 671B-class model: even Unsloth's famous 1-bit dynamic quant is 140 GB, and the 4-bit file is 404 GB. It was never a desktop model, and DeepSeek has since released DeepSeek-V4-Flash under the same MIT licence.

Hardware: What Each Model Needs to Run on Your Own Machine

The rule that explains almost everything about local speed: generating one token means reading the active weights from memory once, so single-user speed is roughly memory bandwidth divided by the weights read per token. This is why a mixture-of-experts model can be faster than a smaller dense one, and why "it fits" and "it is fast" are different questions.

why size misleadsMemory a model needs versus weights read for each token (4-bit)
Qwen3.6 27B (dense)
16.8 GB
16.8 GB
Qwen3-Coder 30B A3B
18.6 GB
1.8 GB
Gemma 4 26B A4B
17 GB
2.1 GB
Qwen3.6 35B A3B
22.4 GB
1.7 GB
Qwen3-Coder-Next
48.5 GB
1.7 GB
Qwen3.5 9B (dense)
5.7 GB
5.7 GB
Unsloth GGUF file sizes from the Hugging Face API (5 October 2026); active parameter counts from each model's config.json and model card, at ~0.56 GB per billion parameters for 4-bit.
Model 4-bit file 8-bit file Minimum NVIDIA GPU Minimum Mac
Qwen3.5 4B 2.7 GB 4.5 GB Any 6 GB card, or CPU only Any 8 GB Mac
Gemma 4 E4B 5.0 GB 8.3 GB 8 GB (RTX 3060 or better) 16 GB Mac
Qwen3.5 9B 5.7 GB 9.5 GB 8 GB 16 GB Mac
Qwen3.6 / 3.8 27B 16.5 to 16.8 GB 28.6 to 29.0 GB 24 GB (RTX 4090, 3090) 24 GB Mac, modest context
Gemma 4 26B A4B 17.0 GB (14.3 QAT) 27.3 GB 24 GB 24 GB Mac
Qwen3 Coder 30B A3B 18.6 GB 32.5 GB 24 GB 32 GB Mac
Qwen3.6 35B A3B 22.4 GB 36.9 GB 32 GB (RTX 5090) 32 GB Mac
Qwen3 Coder Next 48.5 GB 84.8 GB 64 GB+ workstation, or 2x 32 GB 96 GB Mac
DeepSeek-R1 404 GB (140 GB at 1-bit) 713 GB Multi-GPU server 192 GB+ Mac at 1-bit

File sizes are Unsloth's published GGUFs, read from the Hugging Face API on 5 October 2026. Add memory for context on top, and about 2 GB more for MTP builds.

Apple Silicon is a capacity play, not a speed play. A Mac shares one pool of memory between CPU and GPU, so a 64 GB or 96 GB Mac can hold models no consumer graphics card can. But bandwidth is lower than a high-end GPU's, so a dense 27B model generates noticeably slower on most Macs than on an RTX 4090. For mixture-of-experts models the gap narrows, because so little is read per token. Our free models hardware guide lists the bandwidth of every Apple chip and consumer GPU.

What It Costs to Run These Models in the Cloud

Most businesses will not buy a GPU, so the question that matters is the rented one. Start with the price you have to beat: the same models are available as hosted APIs, billed per token.

the price to beatHosted price per million output tokens, October 2026
Llama 3.1 8B
$0
Qwen3.5 9B
$0
Qwen3-Coder 30B A3B
$0
Gemma 4 26B A4B
$0
Qwen3-Coder-Next
$0
Qwen3.6 35B A3B
$0
DeepSeek-R1
$0
Qwen3.8 27B
$0
Qwen3.6 27B
$0
OpenRouter public models API, default listed price per model, read 5 October 2026. Prices move often; check the live listing before relying on any of them.

Two things stand out. The cheapest models to call are the small and sparse ones, and the dense Qwen3.6 27B is the most expensive model in the set at $3.25 per million output tokens, more than DeepSeek-R1. That is not arbitrary pricing. A dense 27B model reads all 16.8 GB of its 4-bit weights for every token, while Gemma 4 26B A4B reads about 2 GB, and providers price the difference.

Renting a GPU for one user usually loses to the API

Now the self-hosted side. Our directory keeps a catalogue of named cloud instances with their on-demand rates (read in September 2026) and memory bandwidth, and every model page ranks them by cost per million tokens. For Qwen3.6 27B, serving one request at a time:

renting a GPU for one userQwen3.6 27B: cost per million output tokens, single user
Hetzner GEX45 (monthly, RTX PRO 4000)
$0
RunPod A100 PCIe
$0
Hosted API (OpenRouter)
$0
RunPod L40S
$0
AWS g5.xlarge (A10G)
$0
Azure NC24ads A100 v4
$0
AWS g6.xlarge (L4)
$0
Our free-models instance catalogue (on-demand rates read September 2026) and the bandwidth method used on every model page: speed is roughly bandwidth divided by the 16.8 GB of 4-bit weights. Ceiling estimates for one request at a time; batching changes the picture completely.
Instance GPU Bandwidth Rate Ceiling speed Cost per 1M tokens
Hetzner GEX45 RTX PRO 4000 Blackwell, 24 GB 672 GB/s about $0.29/hr (monthly) ~40 tok/s $2.01
RunPod A100 PCIe A100 80 GB 1,935 GB/s $1.19/hr ~115 tok/s $2.87
RunPod L40S L40S 48 GB 864 GB/s $0.79/hr ~51 tok/s $4.27
AWS g5.xlarge A10G 24 GB 600 GB/s $1.006/hr ~36 tok/s $7.82
Azure NC24ads A100 v4 A100 80 GB 2,039 GB/s $3.673/hr ~121 tok/s $8.41
AWS g6.xlarge L4 24 GB 300 GB/s $0.805/hr ~18 tok/s $12.52

Speeds are ceilings from the bandwidth method above, for one request at a time. Rates are indicative snapshots from September 2026; follow each provider's link for the live price.

Only two instances beat the hosted API for a single user, and the hyperscalers cost 2.4x to 4.1x the API price per token. For smaller models the gap is far wider. Qwen3.5 9B is $0.15 per million tokens hosted, and the cheapest single-user rental in our catalogue works out at about $0.68, 4.6 times the API. The slow-but-cheap instances are the worst value of all: AWS's L4 is the cheapest AWS GPU by the hour and among the most expensive per token, because at 300 GB/s it bills for longer to produce the same output.

When self-hosting does pay

Self-hosting wins in three situations, and they are worth naming precisely:

  1. Concurrency. The figures above serve one request at a time. Serving engines such as vLLM batch many requests through the same weights, which multiplies throughput on the same rented card. Unsloth's NVFP4 tests measured Qwen3.6 35B A3B at up to 17,561 tokens per second on one B200 at high concurrency. A shared endpoint for a team or a fleet of agents is a different economic animal from one chat window.
  2. A flat monthly box. Hetzner bills its GPU servers monthly, so a GEX45 that runs all month costs the same whether it serves one request or a million. Once utilisation is high, the per-token cost falls below any metered API.
  3. Data that cannot leave. If the documents are patient records, legal files or financial data, the comparison is not with the API at all, because the API is not an option. More on that below.

Three costs never appear in a headline hourly rate: egress when results leave the provider, storage for the weights (16.8 GB pulled on every cold start for the 27B, 48.5 GB for Qwen3 Coder Next), and idle time, because an hourly instance bills while it waits for the next request. Our guide to self-hosted LLMs versus cloud APIs works through the break-even point in more detail, and the cost of AI agents covers the build side of the budget.

Licences and Data Residency

Twelve of the fifteen repositories in the verified top 15 are Apache 2.0, which permits commercial use, modification and redistribution with no user thresholds. DeepSeek-R1 is MIT, equally permissive. The two Llama 3.1 rows are the exception: commercial use is allowed below 700 million monthly active users, derivatives must carry "Llama" in the name, and the repository needs manual access approval, which is a real obstacle for any pipeline that re-downloads weights unattended.

Most of the models here come from Alibaba's Qwen team. A common worry is that a Chinese-developed model sends data to China. Self-hosting removes that concern in a way no vendor contract can. When the weights run on a server you control, there is no model vendor in the data path at all, so the only jurisdiction question left is where your own server sits. We put it to clients this way: self-hosting turns "which AI vendor's terms cover us" into "which country is our box in", a decision made once at setup. If you rent the server, you still need a data processing agreement with the hosting provider, but that is one agreement with one infrastructure company. Your own procurement policy may still restrict model origin, and that is a policy question, not a technical one.

How to Choose Among the Most Downloaded Unsloth Models

Choose by the job first, then by the hardware you have, then by format. The decision tree below summarises the recommendations above.

picking a modelWhich of the most downloaded Unsloth models to run

For coding agents, use Qwen3.8 27B on a 24 GB card, or Qwen3 Coder Next with 64 GB or more. For multilingual chat and image work, use Gemma 4 26B A4B. For routing, extraction and tool calls, use Qwen3.5 9B, or Qwen3.5 4B on a CPU. For audio on the device, use Gemma 4 E4B. For fast autocomplete, Qwen3 Coder 30B A3B. If a single user is the only load and the data may leave, the hosted API is cheaper than renting a GPU.

The recommendations in this guide.

Then apply the economic filter. If one person will use the model and the data is allowed to leave the building, call the hosted API: it is cheaper per token than renting for almost every model here. If several people or agents will share it, or the data must stay on infrastructure you control, self-host, and pick the instance on cost per token rather than hourly rate.

Whatever you pick, test it on your own work before you commit. Language models are not deterministic: even with temperature at zero and structured output switched on, the same input can come back differently across model versions, and a quantised file is a slightly different model from the full-precision one. A swap that looks like a free upgrade can return perfectly valid output that a downstream step no longer handles, and nothing fails loudly. Keep a small set of real inputs with known-good outputs, run every candidate model and every candidate quant through it, and keep logging in production; our guide to monitoring AI in production covers what to watch.

The Competitor Pulse Check

Factor ValueStreamAI Approach Typical "Best Local LLM" Article
Model list Re-verified against the Hugging Face API on the day, with the date Copied from a chart or from memory; often lists models that have been superseded
Download figures Explains what a download is (any GET or HEAD request) and shows the 30-day trend Treats all-time downloads as a popularity vote
Hardware sizing Real file sizes for each quantisation, plus the bandwidth rule for speed Rounded "needs about 16 GB" figures
Cloud cost Named instances ranked by cost per million tokens against the hosted API Hourly rates, or no cloud costs at all
Formats Explains when NVFP4, MTP, QAT and bnb-4bit actually help, including where headline speed-ups were measured Lists formats without saying what hardware they need
Licences Commercial verdict per model, including gating "Open source" applied to everything

What We Would Actually Deploy

A model on its own is not a system, and naming the software matters more than naming the model. For a single workstation, Ollama or LM Studio loads a GGUF in minutes, and Open WebUI gives non-technical staff a ChatGPT-style screen pointed at your own endpoint. Ollama ships official builds of Qwen3.5, Qwen3.6, Qwen3.8, Gemma 4 and Qwen3 Coder Next, so most teams never need to touch a raw file. Once more than a handful of people or agents query the model at the same time, move to vLLM or SGLang, which batch requests and are the only way to run NVFP4. llama.cpp sits underneath most of the desktop tools and now runs the MTP builds directly.

The second thing the download chart cannot show is that a better model does not make an AI system know your business. The most common misunderstanding we meet is a team that has access to a capable model and expects it to answer questions about their own operations. The model is a capable reader that has never been handed the filing cabinet, and swapping in a stronger reader changes nothing. Getting your documents and records out of the CRM, the shared drive and the accounting system, cleaned up and indexed where the model can reach them, is the larger part of every successful deployment, and it is the same work whether the model is Qwen, Gemma or a hosted API.

If you want help choosing a model, sizing the hardware or building the data layer around it, our AI consulting team does exactly this, and our AI agent development service builds the agents that run on top. For a quick price, the automation quote tool gives an estimate in a few minutes, and our pricing page lists the engagement tiers. Teams in healthcare should also read private AI for medical practices, which applies the same self-hosting approach to patient data.

Frequently Asked Questions

What are the most downloaded Unsloth models?

As of 5 October 2026, the most downloaded Unsloth repository is Qwen3-Coder-30B-A3B-Instruct-GGUF with 28.6 million all-time downloads, followed by Qwen3.8-27B-GGUF at 16.9 million and gemma-4-26B-A4B-it-GGUF at 11.2 million. The widely shared chart of the top 15 leaves out the Qwen3 Coder model entirely. Its fifteen rows cover eleven distinct models, because several are the same model in different file formats.

Do Hugging Face download counts show how many people use a model?

No. Hugging Face counts every GET or HEAD request to a model's counted files as a download, counts each GGUF file in a repository, and does not exclude automated pipelines. A fine-tuning notebook that loads a model on every run adds a download each time. The 30-day download figure is a better guide to current demand than the all-time total.

What hardware do I need to run Qwen3.6 27B or Qwen3.8 27B locally?

A 24 GB GPU such as an RTX 4090 or RTX 3090 runs either model at 4-bit, since Unsloth's 4-bit GGUF files are about 16.5 to 16.8 GB. A 24 GB Mac also works with modest context, and a 32 GB Mac gives more room. Add about 2 GB if you use the MTP build for faster generation.

What is the difference between GGUF and NVFP4?

GGUF is the llama.cpp file format that runs almost anywhere: CPUs, Apple Silicon, and NVIDIA or AMD GPUs, through tools such as Ollama and LM Studio. NVFP4 is NVIDIA's 4-bit floating-point format that runs only on Blackwell GPUs, through vLLM or SGLang. NVFP4's large speed-ups appear when serving many users at once; for a single user, Unsloth measured its Qwen3.6 27B NVFP4 build at only about 1.03 times the decode speed of other NVFP4 builds.

Is it cheaper to self-host an open model or use an API?

For a single user, the API is usually cheaper. Qwen3.6 27B costs $3.25 per million output tokens hosted, and only two instances in our catalogue beat that one request at a time, while AWS, Azure and Google Cloud GPUs cost 2.4 to 4.1 times more per token. Self-hosting pays when many users or agents share the model, when you rent a flat-rate monthly server and keep it busy, or when the data is not allowed to leave infrastructure you control.

Can I use these models commercially?

Twelve of the fifteen most downloaded Unsloth repositories are under Apache 2.0, which allows commercial use with no conditions, and DeepSeek-R1 is under the equally permissive MIT licence. The exception is Llama 3.1 8B, whose community licence allows commercial use below 700 million monthly active users, requires "Llama" in derivative names, and sits behind a manual access approval.

Which model in the list is best for coding?

On a 24 GB card, Qwen3.8 27B is the strongest choice in the list, scoring 61.7 on SWE-bench Pro on Qwen's own table. With 64 GB or more of memory, Qwen3 Coder Next is fast and capable at 74.2% on SWE-bench Verified. Qwen3 Coder 30B A3B is the fastest for autocomplete on a 24 GB card but is clearly behind both on autonomous coding tasks.

What's Next

The download chart is a good map of what people are trying, and a poor guide to what you should run. Start from the job, check the memory you have against the real file sizes, and compare self-hosting with the hosted price on cost per token, not per hour. Every model covered here has a page in our free AI models directory with its licence verdict, hardware fit and a ranked table of cloud instances. When you are ready to put one into production, talk to us about the deployment and the data layer around it.

Disclaimer: This article is for informational purposes only and does not constitute financial, legal, or professional advice. Consult a qualified professional before making business or investment decisions.
ShareLinkedInX / Twitter
MK
Co-founder · AI & Automation Engineering

Muhammad Kashif is co-founder of ValueStreamAI, leading technical delivery and AI strategy. He designs and ships custom agentic AI and healthcare automation systems for clients across the US and UK. More about Muhammad Kashif →

← back to blog
LIMITED PILOT SLOTS EACH MONTH

Thirty minutes.
We'll tell you exactly
where your ROI is.

No sales deck. No 50-page report you have to pay for before anything gets built. Just a direct conversation about which of your workflows are costing the most and whether AI can fix them. If there's no compelling answer, we'll say so. And it's a conversation with Kash, our founder, not a rep reading from a script, because the person who built this business is the one who should understand yours.

Book a strategy call ->
info@valuestreamai.com - operating across US + UK