Harness engineering is the work of designing everything around an AI model that turns it into a reliable agent: the loop it runs in, the tools it can call, the context it sees, what it is allowed to touch, and the checks that catch its mistakes. It matters more than most buyers assume. On the public Terminal-Bench 2.1 leaderboard, the same model scores up to 8.1 points differently depending on which harness runs it, and the vendor's own harness is not always the best one.
That changes the question a business should ask. Choosing between GPT, Claude, Qwen or DeepSeek is the visible decision. Choosing and configuring the harness decides whether the agent finishes the job, how much it costs to run, and whether your data ever leaves your network. This guide covers how agent harnesses work with both open-source and closed models, how skills and MCP fit in, how to build a private one, and what it costs.
| Metric | October 2026 figure |
|---|---|
| Widest score gap for one model across two harnesses | 8.1 points (Gemini 3 Pro, Terminal-Bench 2.1) |
| Model providers one open-source harness supports | 75+ (OpenCode, including local models) |
| Public MCP servers when MCP moved to a neutral foundation | 10,000+ (Anthropic, December 2025) |
| Code written by agents in OpenAI's harness engineering experiment | About 1 million lines, none by hand (OpenAI, via InfoQ) |
| Stars on DeepSeek's own open-source harness, two months after launch | 245,000+ (GitHub, October 2026) |
What Is Harness Engineering?
Harness engineering is the discipline of designing the runtime around a model rather than waiting for a better model. The shortest definition comes from Thoughtworks' Birgitta Böckeler, writing on martinfowler.com in April 2026: Agent = Model + Harness, where the harness is everything except the model.
The term was popularised in February 2026 by OpenAI's post "Harness engineering: leveraging Codex in an agent-first world", which described a five-month internal experiment in which, as InfoQ reported, a small team shipped roughly a million lines of code with no manually written source. The lesson OpenAI drew was not about the model. When the agent struggled, the fix was to give it what it lacked: a structured documentation directory that acts as the single source of truth for agents, architecture rules enforced mechanically by linters and CI, and telemetry the agent could read to reproduce its own bugs.
Harness engineering is the third step in a sequence most teams have already lived through. Prompt engineering is about how you phrase the request. Context engineering is about what information reaches the model. Harness engineering is about how the agent runs: its loop, its tools, its limits and its feedback.
For a business, the practical translation is simple. Most of the reliability, cost and privacy of an AI agent is decided by the harness you put around it, and the harness is the part you control. You cannot retrain GPT-5.5. You can decide exactly what it is allowed to do.
Why the Harness Matters More Than the Model You Pick
The harness can move an agent's results by more than the gap between competing models. Terminal-Bench 2.1, led by Stanford and the Laude Institute and hosted by Snorkel AI, tests agents on real terminal tasks and publishes the same model under different harnesses. Its own FAQ states that "the same model can therefore receive different Terminal-Bench 2.1 scores when paired with different agent harnesses".
| Model | Vendor's own harness | Terminus 2 harness | Gap |
|---|---|---|---|
| Claude Fable 5 | 83.8% (Claude Code) | 80.4% | 3.4 points |
| GPT-5.5 | 83.1% (Codex CLI) | 78.0% | 5.1 points |
| Claude Opus 4.7 | 68.9% (Claude Code) | 66.1% | 2.8 points |
| Gemini 3 Pro | 65.8% (Gemini CLI) | 73.9% | 8.1 points, the other way |
| Gemini 3.1 Pro | 65.8% (Gemini CLI) | 65.6% | 0.2 points |
Two findings in that table matter for anyone buying or building an agent.
First, a five-point harness gap is as large as most model upgrades. GPT-5.5 in Codex CLI beat GPT-5.5 in Terminus 2 by 5.1 points. Teams routinely pay for a more expensive model to gain less than that.
Second, the vendor's harness is not automatically the best one. Gemini 3 Pro scored 8.1 points higher in the independent Terminus 2 harness than in Google's own Gemini CLI. A harness is engineering, and engineering can be done better or worse by anyone.
The honest caveat: a leaderboard is not a controlled experiment. Submissions differ in prompts, versions and time budgets as well as harness, and a 2026 survey of agent and harness design makes exactly that point. The direction is still consistent across sources. Harness choice moves results by several points for the same model.
This is also why so many products sold as "AI agents" disappoint. We see the same signature again and again in proposals clients bring us: a database, an API key for one closed model, documents embedded into a vector store, and a thin interface on top. That is a model with almost no harness. It can answer questions about documents. It cannot take an action, check its own work, or recover from a mistake, because none of that machinery exists. If you want the longer version of how to spot it, our guide on how to tell whether an AI agency can actually build it walks through the tells.
The Six Parts of an AI Agent Harness
Every agent harness, whether it is Claude Code, OpenCode or one you build yourself, is made of the same six parts. Knowing them is what lets you compare harnesses on substance rather than on demos.
1. The loop. The agent reads its context, decides on an action, takes it, observes the result and repeats until the task is done or it stops. Everything else hangs off this loop.
2. Tools. Single actions the model can call with structured arguments: read a file, run a command, query a database, send an email. A model with no tools can only talk.
3. Context management. What the model sees on each turn: instructions, files, retrieved documents, the history so far. Context windows are finite, so the harness decides what to keep, summarise or drop.
4. Permissions and sandboxing. What the agent is allowed to touch, and where it runs. A good harness asks before destructive actions, scopes credentials, and runs the agent somewhere a mistake cannot reach production.
5. Sensors. Checks that run after the agent acts: tests, linters, schema validation, business rules, and sometimes a second model reviewing the output. Sensors are how an agent learns it was wrong.
6. State and memory. What persists between steps and between sessions: the task log, decisions made, facts about the user or the business.
The harness loads guides such as instructions and skills before the agent acts. The model picks an action, the harness checks it against permissions, runs the tool in a sandbox, then sensors such as tests and rule checks inspect the result. A failed check goes back into the loop as feedback; a passing result either continues the task or ends it.
The part teams skip is the fourth. We have recovered clients from production incidents where an agent wrote bad data into a live CRM, and one where an agent sent real payment notifications during what was meant to be a test. Neither needed a better model. Both needed a sandbox that mirrored production, a staging CRM and mocked payment and messaging tools, before the agent touched anything real. Docker, a staging server or a cloud sandbox all work. The discipline matters more than the technology.
Guides and Sensors: How Harness Engineering Makes Agents Reliable
The most useful framework for reliability comes from Böckeler's work: a harness controls an agent with guides that act before it does something, and sensors that observe after it acts. Each can be computational, meaning deterministic code, or inferential, meaning another model's judgement.
| Control | Before the agent acts (guides) | After the agent acts (sensors) |
|---|---|---|
| Computational (fast, deterministic) | Code templates, type definitions, language servers | Tests, linters, type checks, schema validation, structural rules |
| Inferential (flexible, less predictable) | Instructions files, skills, retrieved documentation | AI review, an LLM acting as judge |
The rule that falls out of this is the one most failed agent projects break: you need both sides. Böckeler notes that an agent with only feedback repeats its mistakes, and an agent with only rules never learns whether the rules worked.
Prefer computational controls wherever you can. They run on every change, cost nothing per call, and never have an off day. Large language models are non-deterministic even with temperature at zero and structured output enforced: an agent can return perfectly valid JSON with the wrong answer inside it, or call the right tool with a parameter that is technically valid and contextually wrong. A schema check catches the first kind of error. Only a business rule written in code catches the second reliably. That is why every agent we put into production validates its decision before the tool executes, not after.
Böckeler is also candid about what a harness cannot yet do. Checking that the behaviour is actually correct, rather than merely well-formed, remains "the elephant in the room", and AI-written tests are not yet trustworthy enough to replace human supervision. A harness directs human attention to where it matters. It does not remove the need for it.
Skills, MCP and Tools: What Each Layer of a Harness Does
Three terms get used interchangeably and should not be. They solve different problems, and a well-built harness uses all three.
| Layer | What it is | What it solves | Example |
|---|---|---|---|
| Tools | Single functions the model calls | Doing one action | read_file, run_sql, send_email |
| MCP servers | A protocol connecting the harness to outside systems | Access to live data and other software | A CRM, a database, GitHub |
| Skills | Folders of instructions, scripts and references loaded on demand | Know-how: how your business does a task | "How we prepare a client onboarding pack" |
Tools are the verbs. MCP is the plumbing. The Model Context Protocol lets any compatible harness talk to any compatible system without custom integration code for each pair. Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded with Block and OpenAI, in December 2025, when there were already more than 10,000 public MCP servers; our explainer on the Agentic AI Foundation covers what that governance change means.
Skills are the know-how. A skill is a folder containing a SKILL.md file, with a short name and description at the top and instructions underneath, plus optional scripts and reference files. The harness shows the model only each skill's short description until a task needs it, then loads the rest. That progressive loading is why an agent can carry dozens of skills without filling its context window. Anthropic introduced skills in October 2025 and published them as an open standard that December, and harnesses well beyond Anthropic's now load them, including DeepSeek's own harness and Pi. The largest public collection is anthropics/skills; check each skill's own licence, because the repository has no single one. We also keep free skill templates you can adapt.
For a business, skills are the most underrated part of the stack. They are how your process, your formats and your rules get into the agent without retraining anything, and they are plain files your team can read, review and version.
One security rule covers all three layers: anything the agent reads from outside is untrusted input, never an instruction. An agent that reads email, documents or the output of a third-party MCP server inherits hidden prompt injection: instructions planted in white-on-white text or invisible markup that a human never sees and a model obeys. The "ShadowLeak" case in September 2025 used exactly that to pull data out of an AI email agent. The harness has to enforce the boundary, through permissions that limit what a hijacked agent could do, and sensors that check actions against rules rather than trusting the model's reasoning. Install MCP servers and skills from sources you would trust with the same access as an employee, because that is the access they get.
Harness Engineering With Open-Source Models: Building a Private AI Agent
Open-weight models and open-source harnesses together give you a fully private agent: nothing leaves hardware you control. The harness does not care whether the model behind it is closed or open, as long as the model speaks an API the harness understands, and almost every local runtime now exposes an OpenAI-compatible one.
Here is the concrete version. OpenCode, the most widely used open-source alternative to Claude Code, supports more than 75 model providers and documents local models directly. Pointing it at a model running in Ollama on your own machine is one configuration block:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"ollama": {
"npm": "@ai-sdk/openai-compatible",
"name": "Ollama (local)",
"options": { "baseURL": "http://localhost:11434/v1" },
"models": { "qwen3-coder:30b": { "name": "Qwen3 Coder 30B" } }
}
}
}
The same pattern works for LM Studio and llama.cpp, which OpenCode's documentation lists alongside Ollama. Two things decide whether it actually works in practice:
- The model must be good at tool calling. An agent is mostly tool calls. OpenCode's own documentation recommends Qwen-Coder or DeepSeek-Coder variants for strong local tool calling, and Ollama documents tool calling with Qwen3 models.
- Give the model enough context. OpenCode's documentation notes that when tool calls fail on Ollama, the fix is usually raising
num_ctxto between 16K and 32K. Small default context windows silently break agents, because the tool definitions alone can fill them.
On hardware, our free models directory records measured memory needs for each model. Three practical anchors, all Apache 2.0 licensed:
| Model | Memory at 4-bit | Runs on your own hardware | Or rent (indicative rates) |
|---|---|---|---|
| gpt-oss-20b | 13 GB | A 16 GB GPU or Mac | AWS g6.xlarge (L4, 24 GB) at about $0.81/hr, or a Hetzner GEX45 at about EUR 214 a month |
| Qwen3 Coder 30B A3B | 18.6 GB | A 24 GB GPU or 32 GB Mac | AWS g5.xlarge (A10G, 24 GB) at about $1.01/hr, or RunPod L40S (48 GB) at about $0.79/hr |
| gpt-oss-120b | 65 GB | An 80 GB GPU, or 96 GB+ of unified memory | RunPod A100 80 GB at about $1.19/hr, or RunPod H100 at about $1.99/hr |
Most businesses rent rather than buy, and where you rent matters more than which model you pick. For identical silicon, the big three clouds charge roughly 2 to 5 times what specialist GPU clouds do: an A100 80 GB is about $3.43 per GPU-hour on AWS against $1.19 to $1.39 on RunPod. The g5.xlarge above also shows why memory bandwidth matters more than memory size for an agent: it has the same 24 GB as the g6.xlarge but twice the bandwidth, so it generates tokens roughly twice as fast for about 25% more per hour. These are indicative September 2026 rates; check the AWS, RunPod and Hetzner pricing pages before you budget, and add egress, storage for the model weights and idle time, which no headline rate includes.
Name the software when you plan this, not the category. "A private AI agent" is not a plan. "Ollama serves Qwen3 Coder on our own GPU, OpenCode is the harness, and it can only reach our staging database" is a plan someone can build and audit. Our guide to self-hosted AI versus cloud APIs covers the serving layer in more depth.
Self-hosting also changes the compliance question entirely. When the model runs on a server you control, there is no AI vendor's contract to evaluate. A UK or EU business running an open model on an EU-hosted server keeps data inside the UK or EU end to end; a US healthcare practice running the same model on US hardware it controls does the same for HIPAA. Self-hosting means choosing where your own server lives, not checking whether a vendor's terms happen to cover your jurisdiction. One warning from the open-harness world: DeepSeek's own harness sends requests to DeepSeek's China-hosted API by default, so point it at a model you host before any client data touches it.
Closed Models, Open Harness: The Hybrid Most Businesses Should Start With
Most businesses should not start fully self-hosted. The pragmatic design is an open-source harness with a model router: a closed frontier model for the hardest reasoning, and an open model you host for anything touching sensitive data or running at high volume.
The harness inspects each task. Work involving sensitive or regulated data goes to an open model on your own infrastructure. Other high-volume routine work also goes to the self-hosted model to control cost. The remaining hard reasoning goes to a closed frontier model through its API.
This is where owning the harness pays off. Because the harness, not the model, holds your tools, skills, permissions and sensors, you can move a workload from a closed model to an open one by changing configuration rather than rebuilding. That is the strongest protection against vendor lock-in available today, and it is the main reason we build on open harnesses even when the client's first model is a closed one.
It also keeps costs honest. Model choice per step matters as much as model choice overall: for bounded routing and yes-or-no decisions, specialised decision models now cost a fraction of a frontier model per call, a pattern we covered in our analysis of Jev and Needle 3.
How to Build Your Own AI Agent Harness: A Six-Step Plan
Do not write a harness from scratch. Start from an open one and engineer the parts that are specific to your business. This is the plan we follow.
- 01Pick one job and its definition of done
One workflow, one measurable outcome, agreed before any build.
- 02Audit what the agent must touch
Every system it reads or writes, with API access confirmed.
- 03Choose the model route
Closed API, open weights you host, or a hybrid split by data sensitivity.
- 04Start from an open harness
Configure, do not rebuild: models, tools, permissions, sandbox.
- 05Write the guides and the sensors
Skills and instructions before it acts, tests and checks after.
- 06Shadow run, then widen
Run alongside the people doing the job, then extend permissions.
1. Pick one job and its definition of done. One workflow, one measurable outcome, agreed by the people who own it before anything is built. An agent without a definition of done cannot have sensors, because there is nothing to check against.
2. Audit everything the agent must touch. List every system it reads from or writes to and confirm each one has a documented, accessible API, and that someone can hand over the credentials. Discovering in week seven that a core system has no API is a re-architecture. Discovering it in week one is a conversation. It is also where most "the AI does not know our business" problems are really found: the agent has no access to the data, which is a data pipeline problem, not a harness one.
3. Choose the model route. Closed API, open weights you host, or the hybrid above, decided by data sensitivity, volume and the accuracy the job needs.
4. Start from an open harness and configure it. Pick the harness whose shape fits the job (see the next section), connect your model, and set permissions and the sandbox first, before any tools.
5. Write the guides and the sensors. Guides are your instructions file and skills. Sensors are tests, schema checks and business rules in code, plus AI review only where code cannot judge. Every time the agent fails in testing, add the missing guide or sensor back into the harness rather than patching the prompt. That loop is the core of harness engineering.
6. Shadow run, then widen. Run the agent alongside the people doing the job, compare outcomes, and extend its permissions only as the evidence supports it. We delivered our QA automation engagement exactly this way, embedded in the client's team and CI, and its AI only ever ran on the failure path, which is what kept it both reliable and cheap.
Two warnings from production. Watch the parts you did not change: in voice platforms especially, one configuration change can break behaviour elsewhere that nobody touched, so every harness change needs the sensors re-run, not just the new feature tested. And log everything. Every model call, tool call and decision, with full context, is the only way to debug a non-deterministic system; our guides to AI monitoring in production and AI error handling patterns cover the observability and fallback layers in detail.
The Open-Source Agent Harnesses Worth Knowing in 2026
These are the open-source harnesses we would consider for a client build in October 2026, checked against each repository. Our free models page keeps the full, maintained list with licences.
| Harness | Maintainer | Licence | Best fit |
|---|---|---|---|
| OpenCode | Anomaly (formerly SST) | MIT | A Claude Code style terminal agent with your own choice of model, local ones included |
| Pi | Mario Zechner, now Earendil Works | MIT | A deliberately small core you extend with skills and packages; embeddable via SDK |
| DeepSeek Harness | DeepSeek | MIT | A plugin-everything base for building your own internal agent; still a developer preview |
| Codex CLI | OpenAI | Apache 2.0 | Teams standardised on OpenAI models |
| Goose | aaif-goose (originally Block) | Apache 2.0 | A general desktop agent wired into business tools through MCP |
| OpenHands | OpenHands | MIT | Autonomous agents inside sandboxed containers, self-hosted |
| smolagents | Hugging Face | Apache 2.0 | A small Python library for building a custom business agent |
Read the caveats before choosing. Pi has no built-in permission system by design and its own documentation recommends running it in a container. DeepSeek's harness states in its own safety notes that it has not had a security audit. OpenCode ships almost daily releases, so pin a version for anything in production. A harness's defaults are someone else's decisions about your risk.
What a Custom AI Agent Harness Costs
The software is free; the engineering around it is the cost. Every harness above is open source, and open-weight models cost nothing to license. What you pay for is the work that makes an agent reliable for your specific job: the integration audit, the permissions and sandbox, the skills that encode your process, the sensors that catch its mistakes, and the shadow testing.
Our published bands for this work, the same ones used across our AI system design series. Projects start at $5,000, and if you would rather keep an agent maintained and extended over time, our pricing page sets out ongoing retainers from $1,500 a month:
| Engagement | Typical scope | Price |
|---|---|---|
| Fixed-price pilot | One bounded workflow on an open harness, shadow-tested | $5,000 to $15,000 |
| Architecture review and blueprint | Harness, model route and data layer designed and costed, 2 to 3 weeks | $8,000 to $15,000 (£6,000 to £12,000) |
| Full build, single agent | Production harness, integrations, sensors, observability, 8 to 14 weeks | $22,000 to $55,000 (£18,000 to £45,000) |
Running costs are the other half, and they depend almost entirely on the model route. A closed model bills per token. A self-hosted model costs the hardware or the rented GPU, which our free models directory prices per named cloud instance. Our automation quote tool gives a ballpark for your own workflow in about a minute, and the hire versus automate calculator shows the break-even month against taking on another person. For the full price picture, see what AI agents actually cost in 2026, and for how we scope and deliver this work, our AI agent development service.
The most expensive mistake is not the harness. It is building a capable agent on top of business data nobody has made reachable. An agent that cannot see your CRM, your files or your accounting system will answer confidently from nothing. Budget for getting the data to the agent before budgeting for a better model.
The Competitor Pulse Check
| Factor | ValueStreamAI approach | Typical agent build |
|---|---|---|
| Starting point | An open-source harness, configured and extended | A thin wrapper around one model's API |
| Model choice | Routed per step: closed, open or decision model | One model for everything |
| Lock-in | Model swappable by configuration; you own the skills and sensors | Rebuild required to change vendor |
| Reliability | Guides before, sensors after, computational checks first | Prompt tweaks after each failure |
| Safety | Sandbox, scoped permissions, untrusted-input rule on every source | Agent runs with whatever access it was given |
| Privacy | Self-hosted option in the jurisdiction you choose | Data sent to the model vendor by default |
| Rollout | Shadow run alongside your team, permissions widened on evidence | Switched on and monitored by complaint |
Frequently Asked Questions
What is harness engineering in AI?
Harness engineering is the practice of designing the software around an AI model that turns it into a reliable agent: the loop it runs in, its tools, the context it sees, its permissions and sandbox, and the checks that catch its mistakes. The term was popularised by OpenAI in February 2026 and is summed up as Agent = Model + Harness.
What is the difference between an AI model and an agent harness?
The model generates the next decision; the harness does everything else. The harness feeds the model its context, offers it tools, decides what it is allowed to do, runs its actions, and checks the results. The same model can score up to 8.1 points differently on Terminal-Bench 2.1 depending on the harness running it.
Can I build an AI agent harness with open-source models?
Yes. Open-source harnesses such as OpenCode, Pi and DeepSeek Harness can drive open-weight models served locally through Ollama, LM Studio or llama.cpp, which keeps every request on hardware you control. Choose a model with strong tool calling, such as a Qwen Coder variant, and give it a context window of at least 16K to 32K tokens.
What is the difference between skills and MCP?
MCP connects an agent to outside systems such as a CRM or database, so it is about access. Skills package know-how: a folder with a SKILL.md file of instructions and optional scripts, loaded only when a task needs it. Most production agents use both, with skills often telling the agent how to use the tools that MCP provides.
Is a private AI agent more secure than one using a closed model?
A private agent keeps your data on infrastructure you control, which removes the AI vendor from the data path and settles where the data lives. It does not make the agent safe by itself. Permissions, sandboxing and treating every external input as untrusted are still required, because prompt injection works the same way against an open model as a closed one.
How much does it cost to build a custom AI agent harness?
The harness software and open-weight models are free; the cost is the engineering. A fixed-price pilot on one workflow typically runs $5,000 to $15,000, an architecture blueprint $8,000 to $15,000, and a full production build of a single agent $22,000 to $55,000, before running costs for the model.
What's Next
If you are weighing whether to build an agent in-house, buy a product, or have one built on an open harness you will own, the deciding questions are the ones in this guide: which model route your data allows, what the agent must touch, and who writes the sensors. The complete guide to building AI agents covers the agent design itself, and our piece on hiring AI agent developers covers what to look for in the people who build it.
To talk through your own workflow, book a free strategy call. We will tell you which harness and model route fit the job, what it would cost, and whether you need us at all.
Muhammad Kashif is co-founder of ValueStreamAI, leading technical delivery and AI strategy. He designs and ships custom agentic AI and healthcare automation systems for clients across the US and UK. More about Muhammad Kashif →
