Jev from TypeSafe AI and Needle 3 from Cactus Compute are two AI models released in mid-September 2026 that share one idea: most steps in business automation are not conversations, so they should not be run by a chat model. Jev returns a typed decision with a probability instead of text, for about $0.042 per million input tokens with free output. Needle 3 turns a command into a function call inside a single 8 to 29 MB file that runs offline on a phone or a Raspberry Pi. Both are genuinely useful. Both also arrive with headline claims that need a condition attached, and this guide sets out which is which.
The short version for a business owner: Jev is a fast, cheap decision layer for workflows that run in the cloud, and Needle 3 is a tiny command layer for devices that cannot or should not call the cloud. Neither replaces a large language model for writing, reading documents or explaining a decision to an auditor.
| Metric | September 2026 figure |
|---|---|
| Jev price | $0.042 per million input tokens, output free (TypeSafe) |
| Jev latency | 70-500 ms per decision (TypeSafe) |
| Jev agreement with reference answers | 67.8%, vs 74.1% for GPT-5.6 Sol and 73.1% for Claude Opus 5 (TypeSafe evals) |
| Needle 3 size | 8-29 MB, one file, runtime engine under 1 MB (Cactus Compute) |
| Needle 3 on Mobile Actions | 86.0%, vs 88.4% for DeepSeek V4 Flash in the cloud (Cactus Compute) |
| Needle 3 licence | Apache 2.0, commercial use allowed (Hugging Face tag) |
Why AI Models That Do Not Chat Matter for Business Automation
AI models that do not chat matter because most automation steps are small, bounded decisions, and running each one through a general chat model is slow, expensive and occasionally malformed. Look at any real workflow we build and the model calls are rarely "write an email". They are "which queue does this ticket belong in", "is this invoice a duplicate", "should this agent's proposed action go ahead", "which tool does this request need".
A frontier language model can answer every one of those. It does it by generating text token by token, which takes seconds, costs output tokens, and then has to be parsed back into something the software can use. TypeSafe's own launch material puts structured-output error rates for frontier models at between 0.58% and 45.5% depending on the task. At 500 runs a day against a live CRM, even the bottom of that range is a daily support ticket.
We have written before about why RPA, AI agents and plain code each own different steps. Jev and Needle 3 add two new categories to that map: a decision model that sits between rules and a full agent, and a tool-calling model small enough to live on the device itself.
It cannot return an answer outside the options you declared. It can still pick the wrong option, confidently.
True, and output is tiny by design. TypeSafe itself says it cannot prove the pricing is not subsidised.
Only after fine-tuning on your tools. Out of the box, DeepSeek wins all six of Cactus's own benchmarks.
Inference is offline, but the binary sends anonymous telemetry by default. Turn it off before a private deployment.
What Is Jev?
Jev is a "System One Model" from TypeSafe AI: you give it the current state of a task and a set of typed questions, and it returns an answer to each question with a calibrated probability, in a single pass, with no text generation at all. TypeSafe was founded by Diogo Almeida, a former OpenAI researcher who worked on the reinforcement learning from human feedback behind ChatGPT, and raised $40 million. Jev entered early access on 15 September 2026.
How Jev works, in plain terms
A normal language model writes its answer one token at a time. Jev is non-autoregressive: it produces every answer in one query. TypeSafe trained it with a method it calls Reinforcement Learning for Calibrated Decisions, and has not published the architecture, the model size or the training data.
You call it with two things:
- State: the situation, as text or structured data. A support ticket, an agent's proposed next action, a transaction record.
- Questions: what you want decided, each one typed. There are three primitives:
- Choice: pick one of a declared list, such as
billing,technical,salesorspam, with up to 255 options. - Score: a number within declared bounds, such as a 0 to 100 lead score.
- Noul: a yes or no, returned as a probability.
- Choice: pick one of a declared list, such as
The response is the typed values plus probabilities. There is nothing to parse, because there is no text.
Where Jev is already available
Jev launched with an unusual amount of ecosystem support for a model that is days old. It is callable through TypeSafe's own API and Python and JavaScript SDKs, through Pydantic AI as typesafe:jev-latest, through a LiteLLM pass-through with cost tracking, and on Cloudflare Workers AI with a documented 32,000-token context window. Early users quoted by TechCrunch reported results 5 to 18 times faster than an OpenAI model at Vercel, and 10 to 20 times cheaper than Gemini at Bryo AI. Demand was high enough that TypeSafe briefly ran out of serving capacity.
What Jev is designed for
TypeSafe positions Jev as a "smart if-statement": the fuzzy judgement inside a workflow that plain code cannot make, but that does not need a paragraph of prose either. Its own recommended uses are classification and routing at volume, scoring, guardrailing the output of other models, and real-time decisions inside an application loop. Its stated anti-patterns are writing code, emails or summaries, and multi-step reasoning that needs an explanation.
Jev's Benchmarks, Read Carefully
Jev's benchmarks show a model that is roughly as accurate as a mid-tier frontier model on bounded decisions, dramatically faster and cheaper, and noticeably less accurate than the best models on harder cases. That is a useful profile. It is not the profile the phrase "cannot hallucinate" suggests.
TypeSafe's own workflow evaluation covered four production-like tasks: security incident response, agent-trace observability, invoice processing and customer service. As reported by DataCamp:
| Model | Agreement with reference | Cost per case | Latency per case |
|---|---|---|---|
| Jev | 67.8% | $0.0004 | 0.4 s |
| GPT-5.6 Terra | 67.9% | $0.0304 | 10.1 s |
| GPT-5.6 Sol | 74.1% | $0.0836 | 23.3 s |
| Claude Opus 5 | 73.1% | $0.1761 | 37.8 s |
Three things in that table matter more than the headline.
First, the accuracy gap is real. Jev matches GPT-5.6 Terra, but trails the strongest models by five to six points overall, and by 17.3 points on invoice processing according to Anthony Maio's independent review. If your workflow is dominated by harder judgements, the cheaper model can cost more in exceptions than it saves in tokens.
Second, the reference answers came from other models. Maio points out that the "correct" labels were produced by averaging two frontier models at high reasoning settings, not by checking real outcomes. That makes the evaluation a measure of agreement, not of being right.
Third, the speed and cost gap is enormous and probably holds. Even allowing for vendor framing, 0.4 seconds against 10 to 38 seconds, and fractions of a cent against several cents per case, is the kind of difference that changes what is economical to automate. Scoring every row in a 50-million-row table stops being a budget conversation.
What "cannot hallucinate" actually means
Jev cannot return an answer outside the options you declared, and it cannot produce malformed output. That is a real and valuable guarantee: no parsing failures, no invented fields. It does nothing to stop Jev choosing the wrong option, misreading the evidence, or attaching a high probability to a bad answer. In Maio's phrase, constraining the shape of the output is not the same as constraining the quality of the judgement.
The Pydantic AI documentation is unusually candid about the weak spots: Jev performs poorly on arithmetic, counting, dates and indirection, is vulnerable to adversarial prompting, and is sensitive to the order in which options are listed. Every one of those shows up in business data. An invoice due-date check is date arithmetic. A "more than two returns in 90 days" rule is counting. Design around them by doing the arithmetic in code and passing the result in as state.
What Jev does not give you
Jev gives no rationale. For many workflows that is fine. For regulated decisions, where an auditor or a customer is entitled to ask why, a bare probability is not an explanation, and you will need either a rules layer that can explain itself or a language model producing the justification alongside. The weights are closed, TypeSafe has not published its calibration method, and it acknowledges it cannot prove its pricing is not subsidised. That is the vendor lock-in question in its newest form: thresholds you tune today are tuned to a model version you do not control.
What Is Needle 3?
Needle 3 is an open-weight model from Cactus Compute, a San Francisco startup from Y Combinator's Summer 2025 batch, that does three narrow jobs on the device itself: tool calling, structured extraction and text embeddings. It ships as a single file of 8 to 29 MB, compressed to about 2.1 bits per weight, with a runtime engine under 1 MB for each platform. It is released under Apache 2.0, so commercial use is allowed, and the GitHub repository passed 12,000 stars within days of the release.
Where Needle came from
Needle has moved fast. The first version, in May 2026, distilled Gemini's tool calling into a 26-million-parameter model, on the argument that tool calling is retrieval and assembly rather than reasoning. Needle 2 followed in the summer at 45 million parameters, a 14 MB file running in 28 MB of RAM and reaching 500 tokens a second on a Raspberry Pi 5. Needle 3 was published to Hugging Face on 16 September 2026.
How Needle 3 works
Needle 3 is 121 million parameters at full depth, but most of them sit in a memory component that is looked up rather than computed, so Cactus says it does the arithmetic of a 50-million-parameter model. It was trained so that every depth from 2 to 20 layers is a usable model: one release can run on a microcontroller at 2 layers and on a laptop at 20.
In use, you describe your app's functions and Needle picks the right ones and fills in the arguments from what the user said. Ask for two things and you get two calls in order. Ask for something no tool covers and you get an empty list rather than a guess. Every response carries a calibrated confidence score, and Cactus's own guidance is a three-band policy: act above 0.7, show the user and ask for confirmation between 0.1 and 0.7, and refuse below 0.1, with a higher bar, such as 0.9, for anything like a payment.
It sees at most five tools per request. With more than five, an embedding lookup picks the five most relevant first. Cactus's examples are all in English, and its own confidence guide notes the score behaves inconsistently in other languages, with Spanish calls measured at 0.0.
Needle 3's Benchmarks, Read Carefully
Needle 3's benchmarks show a model that comes remarkably close to a cloud model on phone-style commands, and falls well behind on everything else until you fine-tune it. Cactus publishes the full chart in its README, which is to its credit, and the chart tells a more careful story than the launch coverage.
| Benchmark | What it tests | DeepSeek V4 Flash (cloud) | Needle 3 (121M) |
|---|---|---|---|
| Mobile Actions | Phone commands, exact call | 88.4% | 86.0% |
| DroidCall | Several calls, in order | 60.5% | 47.0% |
| BFCL v4 | General tool calling | 77.2% | 50.2% |
| DSTC8 | Field extraction (F1) | 80.0% | 40.7% |
| SNIPS gold | Field extraction (F1) | 69.4% | 30.2% |
| SNIPS 7-way | Field extraction (F1) | 66.7% | 24.7% |
The widely repeated claim that Needle beats DeepSeek V4 Flash is true only after fine-tuning. Cactus reports that fine-tuning on the DroidCall dataset lifts every depth of Needle by 18 to 36 points, and that from four layers up the tuned model passes DeepSeek V4 Flash. Untuned, the cloud model wins all six benchmarks.
Extraction is Needle's weak side. On the three extraction tests, the full model scores roughly half of the cloud model, and LFM2.5 1.2B, another small open model, also beats it. If the job is reading invoices or contracts into fields, Needle is the wrong tool.
Where Needle 3 genuinely shines is the first row. 86.0% against 88.4% on phone commands, from a 29 MB file with no network connection, is a real result. For comparison, FunctionGemma 270M, a model more than twice its size, scored 65.1%, and Apple's 3-billion-parameter on-device model scored 57.6%.
Two catches before a private deployment
Telemetry is on by default. The shipped binary sends anonymous usage data unless you set NEEDLE_TELEMETRY=0 and DO_NOT_TRACK=1. Inference itself never touches the network, but "runs offline" and "sends nothing" are different promises, and a healthcare or finance deployment needs the second one.
Fine-tuning trades privacy for calibration. Local fine-tuning keeps your data on your machine but switches the confidence score off. Cactus's hosted fine-tuning keeps the score calibrated on your tools, but uploads your data to its GPUs and is a paid service. For a regulated client that is a real decision, not a footnote.
Where Jev and Needle 3 Fit in Business Process Automation
Jev and Needle 3 fit as two new layers in an automation stack: Jev as the decision layer between rules and agents in cloud workflows, and Needle 3 as the command layer on devices at the edge. The pattern we would build around both is the same one we already use for larger agents: deterministic code stays in control, the model makes the fuzzy call, and a confidence threshold decides whether the system acts, asks, or hands off.
Where Jev earns its place
- Ticket and email routing at volume. Choosing one of eight queues for 40,000 messages a month is Jev's ideal job: bounded answers, high volume, low cost per mistake, and humans downstream to catch errors.
- Lead and risk scoring. A 0 to 100 score with a calibrated probability is exactly what a CRM rule wants to branch on.
- Guardrails on agents. Before an AI agent takes an action in a live system, a fast second model asking "is this action within policy, yes or no" catches a class of failure that prompt instructions alone do not. At 0.4 seconds, the check does not slow the agent down noticeably.
- Loop and failure detection. Reading an agent's trace and deciding "stuck, progressing or failed" is one of TypeSafe's own evaluation workflows.
Where Needle 3 earns its place
- Kiosks and front desks where a spoken request becomes a booking, a check-in or a lookup, and the connection is unreliable.
- Field equipment where a technician logs a reading or triggers a procedure by voice, offline.
- In-app assistants that map "move my 3 pm to Friday" onto your app's own functions without a server round trip.
- Privacy-first deployments where the data should never leave the building. Here Needle fits the argument we make in our self-hosted versus cloud API guide: when the model runs on hardware you control, the compliance question becomes where your hardware sits, not which vendor's contract covers you.
Where neither belongs
- Document extraction. Jev trailed the best models by 17.3 points on invoice processing and Needle scores about half a cloud model on extraction. Our invoice automation guide covers what actually works there.
- Anything that writes. Replies, summaries, reports and code are still language-model work.
- Decisions someone is entitled to have explained. Jev gives no rationale. Needle returns a short
reasoningfield, but it is not an audit trail. - Open-ended, multi-step planning. Both are single-shot by design.
The integration question comes first, as always
A model this cheap does not help if it has nothing to act on. Before choosing between Jev, Needle and a frontier model, check the systems the step touches. If your CRM, booking system and ticketing tool each expose an API or an MCP server, a decision model can be wired in cleanly. If they do not, the project is a browser-automation or data-extraction project first, and the model choice is the easy part. In our experience that single check predicts the cost of an automation project better than any model benchmark.
What Jev and Needle 3 Cost
Jev costs about $0.042 per million input tokens with free output, which worked out at roughly $0.0004 per case in TypeSafe's own evaluation, and Needle 3 is free to use under Apache 2.0 with the cost sitting in the device it runs on. The real cost of either one is the engineering around it.
| Item | Jev | Needle 3 |
|---|---|---|
| Model licence | Closed, API only | Apache 2.0, open weights |
| Usage price | $0.042 per million input tokens, output free | Free |
| Per-case cost (vendor eval) | About $0.0004 | Your own hardware |
| Fine-tuning | Not offered publicly | Local LoRA free; hosted full fine-tune paid |
| Access | Early-access waitlist; also via Cloudflare Workers AI | pip install cactus-needle, Hugging Face |
For a sense of scale: at the per-case costs in TypeSafe's evaluation, 100,000 cases a month would cost around $40 on Jev, against about $3,040 on GPT-5.6 Terra and about $17,610 on Claude Opus 5. Simpler cases, such as routing a short ticket, would cost less on every model. Even if the pricing moves once early access ends, the gap is large enough to change which steps are worth automating.
The engineering around the model is where a budget actually goes: the schema design, shadow testing on your real traffic, threshold calibration, monitoring and the fallback path. A single bounded decision like ticket routing fits inside our fixed-price pilot of $5,000 to $15,000, and our automation quote tool gives a figure for your own workflow in a few minutes.
The Competitor Pulse Check
Most early coverage of Jev and Needle repeats the launch claims. Here is how our approach to evaluating and deploying new models compares.
| Factor | ValueStreamAI approach | Typical launch coverage |
|---|---|---|
| Benchmarks | Reads the vendor's full chart, reports where the model loses | Repeats the single best number |
| "Cannot hallucinate" | Separates output shape from judgement quality | Takes it literally |
| Needle vs DeepSeek | Notes the win requires fine-tuning | Reports the win without the condition |
| Privacy | Flags default telemetry and the fine-tuning trade | Says "runs on-device" |
| Deployment | Shadow mode, thresholds from your data, monitored drift | Swap the model in and measure speed |
| Model choice | Picks per step: code, Jev, Needle or a frontier LLM | One model for everything |
How We Would Pilot Jev or Needle 3
We would pilot either model the same way we introduce any new model into a live workflow: on one bounded decision, in shadow mode first, with thresholds set from your own data rather than the vendor's.
- 01Pick one bounded decision
A step whose possible answers are known in advance and where a wrong answer is cheap to catch.
- 02Run it in shadow mode
The new model decides alongside the current process on real traffic. Nothing it says is acted on yet.
- 03Set thresholds from your own data
Choose the confidence bands for act, confirm and refuse from where it was right and wrong on your cases.
- 04Let it act on the confident band
Automate only above the threshold. Everything else goes to the existing path or a person.
- 05Watch for drift
Re-check calibration monthly and after any model version change. Probabilities that held last month can quietly stop holding.
Start in shadow mode. The new model makes its decision alongside your current process on real traffic, and nothing it says is acted on. Internal testing has a blind spot: the people writing the test cases know what the system is supposed to do. Real users, real tickets and real phrasing surface failure modes that survived weeks of internal checks, and in our deployments the first hundred real interactions almost always surface several.
Calibrate on your own cases. A vendor's calibration was measured on the vendor's data. Plot where the model was right and wrong at each confidence level on your traffic, then set your act, confirm and refuse bands from that. Needle's default of 0.7 is a starting point, not an answer, and Cactus itself suggests a much higher bar for anything that moves money.
Treat every call as untrusted until checked. Language models, and decision models trained like them, are not deterministic in the way code is. The system around them needs input guardrails, output validation before any action, and a log of every decision with the context it saw. That discipline is what makes the difference between a model that saves money and one that quietly writes bad data into a CRM.
Re-check after every model version. Jev is versioned (for example jev-1.13.0 alongside jev-latest). Pin a version in production, and re-run your calibration before moving to a new one.
Frequently Asked Questions
What is the Jev AI model?
Jev is a System One Model from TypeSafe AI, launched in early access on 15 September 2026, that returns typed decisions with calibrated probabilities instead of text. You declare the possible answers as choices, scores or yes-or-no questions, and Jev answers each one in a single pass, typically in 70 to 500 milliseconds.
Can Jev really not hallucinate?
Jev cannot return an answer outside the options you declared or produce malformed output, which removes parsing errors entirely. It can still choose the wrong option or be overconfident, and on TypeSafe's own workflow evaluation it agreed with the reference answers 67.8% of the time, against 74.1% for GPT-5.6 Sol.
How much does Jev cost?
Jev costs about $0.042 per million input tokens, and output tokens are free. In TypeSafe's own evaluation that came to roughly $0.0004 per case, compared with $0.0304 for GPT-5.6 Terra and $0.1761 for Claude Opus 5, although TypeSafe notes it cannot prove the early-access pricing is not subsidised.
What is Needle 3 from Cactus Compute?
Needle 3 is an open-weight AI model from Cactus Compute for on-device tool calling, structured extraction and embeddings, published on 16 September 2026. It ships as a single 8 to 29 MB file under the Apache 2.0 licence, runs offline on phones, single-board computers and microcontrollers, and returns a calibrated confidence score with every call.
Does Needle 3 beat DeepSeek V4 Flash?
Needle 3 beats DeepSeek V4 Flash on tool calling only after fine-tuning on your own tools, according to Cactus Compute. Out of the box, DeepSeek V4 Flash scores higher on all six of Cactus's published benchmarks, narrowly on phone commands at 88.4% against 86.0%, and by roughly double on field extraction.
Is Needle 3 free for commercial use?
Yes, Needle 3 is released under the Apache 2.0 licence, which allows commercial use, modification and redistribution. The shipped binary sends anonymous usage telemetry by default, so set NEEDLE_TELEMETRY=0 and DO_NOT_TRACK=1 before using it anywhere data privacy matters.
Should I use Jev or Needle 3 for business process automation?
Use Jev for fast, high-volume decisions in cloud workflows, such as routing tickets, scoring leads or checking an agent's action against policy. Use Needle 3 for turning commands into tool calls on devices that must work offline or keep data local. Use a frontier language model for document extraction, writing, and any decision that has to be explained.
Can Jev or Needle 3 replace ChatGPT or Claude in my business?
No, Jev and Needle 3 complement general language models rather than replacing them. Neither can write, summarise or explain its reasoning at length, and both are weaker than frontier models on harder judgements and on extraction. They are best used to take the cheap, repetitive decisions off a larger model's plate.
What's Next
Jev and Needle 3 are the clearest sign yet that the AI stack for automation is splitting by job: large models for reading and writing, small or specialised models for deciding and acting. The businesses that benefit first will be the ones that already know which steps in their workflows are bounded decisions and which are genuine judgement.
Needle 3 is now in our free AI models directory, with its licence, benchmarks and the telemetry setting written out. For the wider picture of how these pieces fit together, start with our business automation hub or the guide to AI agents versus chatbots.
If you want to know whether a model like Jev or Needle would cut the cost of a workflow you already run, book a free 60-minute consultation. We will map the steps, tell you which ones are bounded decisions, and give you a fixed price for a shadow-mode pilot before any work begins. See our AI automation solutions for how we scope that work.
Muhammad Kashif is co-founder of ValueStreamAI, leading technical delivery and AI strategy. He designs and ships custom agentic AI and healthcare automation systems for clients across the US and UK. More about Muhammad Kashif →
