homeservicesworkaboutblogfree templatescontactFree Tools →Free AI ModelsResearch LibraryROI CalculatorSavings CalculatorAI Readiness ScoreHire vs. AutomateAutomation Quote
book a 30-min call →
home / free-models / needle-3

Needle 3

Cactus Compute · United States · Apache 2.0

Commercial use: Yes — free for commercial use

Apache 2.0, confirmed on the Hugging Face repository tag. Commercial use, modification and redistribution are all allowed. A separate point worth knowing: the shipped binary sends anonymous usage telemetry by default, so set NEEDLE_TELEMETRY=0 and DO_NOT_TRACK=1 before any private deployment.

What this means for your business

A model about the size of a photo that turns what a user says into the right function call, on the device itself, with no server involved.

Why it should matter to you

Most automation assumes a round trip to a cloud model for every command, which brings latency, a per-call bill and a data-residency question. For narrow, repeatable commands, such as booking a slot, logging a reading or controlling equipment, a model this small can run where the work happens, including places with no reliable connection.

How it connects to our work

We would use it in the same pattern we apply to larger agents: act only above a confidence threshold, confirm in the middle band, refuse below it, and log every call. We would also fine-tune it on the client's own tools before trusting it, because on Cactus's own figures the gap between tuned and untuned Needle is the difference between useful and not.

From our field notesA confidence score is only a guardrail if something in the system acts on it.

Parameters121M (29M to 121M by depth)
Active per token~50M effective
ContextShort prompts, 5 tools per call
Modalitytext, embeddings
Ships as8-29 MB
Format2-bit .cact, about 2.1 bits per weight
LicenceApache 2.0
Last verified2026-09

The runtime engine is under 1 MB per platform, and shallower depths, down to 2 layers, ship as smaller files

Advantages

  • Close to a cloud model on phone commands: 86.0 on Mobile Actions against 88.4 for DeepSeek V4 Flash, in Cactus's own benchmark
  • Runs fully offline once the file is on the device; inference never touches the network
  • Every response carries a calibrated confidence score, so an app can act, ask the user to confirm, or refuse
  • Returns an empty list when no tool covers the request, instead of guessing
  • Any depth from 2 to 20 layers is a usable model, so one release covers a microcontroller and a laptop

Disadvantages

  • Out of the box it trails DeepSeek V4 Flash on all six of Cactus's own published benchmarks
  • Weak at pulling fields out of text: 40.7 F1 on DSTC8 against 80.0 for DeepSeek V4 Flash
  • Sees at most five tools per request; larger tool sets go through an embedding lookup first
  • Fine-tuning locally switches the confidence score off; keeping it means Cactus's hosted fine-tuning, which uploads your data
  • Anonymous telemetry is on by default in the shipped binary

Reach for it when

Turning a spoken or typed command into a function call on a device that cannot, or should not, call the cloud: kiosks, in-app assistants, voice front-ends, field equipment and smart devices.

Where it falls down

Conversation, reasoning and document extraction. If the job is reading invoices or contracts into fields, a larger model does it far better. Without fine-tuning on your own tools, expect accuracy well below a cloud model on anything beyond simple commands.

Running it

Any CPU. A phone, a Raspberry Pi or a microcontroller-class board is enough. See the hardware sizing tables for how that maps to specific chips and cards, and the quantisation guide for what you give up at each bit width.

the model in numbersNeedle 3 at a glance
8-0 MBthe whole model, as one file
0Mparameters at full depth
2-0 layersevery depth is a usable model
under 0 MBruntime engine per platform
Cactus Compute's published figures for Needle 3, as recorded in this directory on 2026-09.
vendor benchmarkMobile Actions: exact tool-call accuracy
DeepSeek V4 Flash (cloud)
0%
Needle 3, 121M
0%
LFM2.5 1.2B
0%
Qwen3.5 0.8B
0%
FunctionGemma 270M
0%
Needle 2, 45M
0%
Apple FM 3B
0%
Cactus Compute's benchmark chart in the Needle GitHub README (961 rows, exact match), read 21 September 2026. Vendor-run and not independently reproduced; DeepSeek V4 Flash was called through its cloud API.

Which cloud instance actually runs this, and what it costs

At Q4_K_M the weights are about 0.03 GB, so 20 catalogued instances fit with room for context, of which the 8 best value are shown. Cheapest per token is Hetzner GEX45 (~$0.00 per million output tokens at an estimated 22400 tok/s). Fastest on a single card is Azure NCads H100 v5 at roughly 111667 tok/s for $0.02 per million tokens, which is the speed-versus-cost trade in one line.

InstanceGPUVRAMBandwidthEst. tok/s$/hr$/M tokens
Hetzner GEX45RTX PRO 4000 Blackwell24 GB672 GB/s~22400$0.290$0.00
RunPod A100 PCIeA100 80GB80 GB1935 GB/s~64500$1.190$0.01
Hetzner GEX131RTX PRO 6000 Blackwell Max-Q96 GB1792 GB/s~59733$1.420$0.01
RunPod L40SL40S48 GB864 GB/s~28800$0.790$0.01
RunPod H100 PCIeH100 PCIe80 GB2000 GB/s~66667$1.990$0.01
Lambda A100 40GBA100 40GB40 GB1555 GB/s~51833$1.990$0.01
Lambda H100 PCIeH100 PCIe80 GB2000 GB/s~66667$3.290$0.01
AWS g5.xlargeA10G24 GB600 GB/s~20000$1.006$0.01

Token generation on a single request is memory-bandwidth-bound, not compute-bound: the GPU must read every active weight from memory for each token it emits. So the ceiling is roughly GPU memory bandwidth divided by the size of the active weights. A 16 GB model on a 300 GB/s L4 tops out near 19 tokens/sec; the same model on a 2039 GB/s A100 tops out near 127. Figures on model pages are that arithmetic, computed from the bandwidth and weight-size columns here. They are a ceiling for one request at batch size 1, before serving overhead, so treat them as an upper bound for comparing instances rather than a benchmark. Batching raises total throughput well above this and does not raise per-request speed.

For identical silicon, the hyperscalers charge roughly 2 to 5 times what the specialist GPU clouds charge. An A100 80GB is about $3.43/GPU-hr on AWS, $3.67 on Azure and $5.03 on Google Cloud, against $1.19-1.39 on RunPod and $1.99 on Lambda. That spread is the single largest cost lever in a self-hosted inference budget, and it is larger than any saving from picking a smaller model.

Three costs that are absent from every headline rate: egress, storage for the weights, and idle time. A 70B model at Q4 is tens of gigabytes that must be pulled to the instance on every cold start, and an endpoint billed hourly is billed while it sits idle waiting for a request. Rates read 2026-09-06; check each provider's live pricing on the directory hub.

Jurisdiction: United States

Best raw capability and the deepest tooling ecosystem. For EU personal data you are relying on a transfer framework rather than on data never leaving the bloc, so check whether your DPA and your customers accept that.

Watch for: Enterprise API tiers usually promise no training on your data; consumer tiers and free tiers frequently do not. The free tier is where this bites.

Frequently asked

Can I use Needle 3 commercially?

Apache 2.0, confirmed on the Hugging Face repository tag. Commercial use, modification and redistribution are all allowed. A separate point worth knowing: the shipped binary sends anonymous usage telemetry by default, so set NEEDLE_TELEMETRY=0 and DO_NOT_TRACK=1 before any private deployment.

What hardware do I need to run Needle 3?

Any CPU. A phone, a Raspberry Pi or a microcontroller-class board is enough. It ships as a single 8-29 MB file (2-bit .cact, about 2.1 bits per weight). The runtime engine is under 1 MB per platform, and shallower depths, down to 2 layers, ship as smaller files.

What licence is Needle 3 released under?

Apache 2.0. Full commercial use, modification and redistribution. Patent grant included. The most permissive licence in common use for open-weight models.

Where can I use Needle 3 for free?

Self-hosting the weights is the free route. Any CPU. A phone, a Raspberry Pi or a microcontroller-class board is enough.

Similar models

Wiring Needle 3 into something real?

We build the evaluation harness, the failover and the cost ceilings around a model like this, so it survives contact with production.

Maps to AI strategy and consulting and AI automation development. Or see it working: our case studies.