The Human Algorithm/What an AI Agent Harness Is, and Why It Decides What Your AI Can Touch
The Human Algorithm • MondayWhat an AI Agent Harness Is, and Why It Decides What Your AI Can Touch
A harness is the program wrapped around an AI model that lets it do things. The model writes text and stops. The harness runs the commands, reads the files, and decides what the model may touch on your computer. Here is what that means for anyone using these tools.
A harness is the program wrapped around an AI model that lets it do things. The model writes text. The harness takes that text, runs the commands in it, reads the files, keeps the notes, and decides what the model is allowed to touch on your computer. When you use Claude Code, or Codex, or Antigravity CLI, the model is the part that thinks and the harness is everything else.
The word turns up constantly now and almost nobody defines it. People say a model is good at coding, or bad at coding, when what they actually tested was a model inside somebody else’s harness.
What a harness actually is
Think of an engine and a car. The engine makes power. On its own it sits on a stand and spins. The car is the steering, the pedals, the brakes, the doors that lock, the seat belt, the fuel line. You do not drive an engine. You drive the thing built around it, and two cars with the same engine are not the same car.
The model is the engine. Words go in, words come out. It cannot open a file, run a command, or remember yesterday.
Anthropic’s own documentation for tool use describes the mechanism plainly, and it is worth knowing because it settles the question. When the model wants to use a tool, it stops and says so, in a structured message. Then, in their words, your code executes the operation and sends back a result. The model never touches the file. Something else does, on its behalf, and that something else is the harness.
The four jobs it does while you are not watching
It runs the loop. A model produces one answer and stops. The harness takes that answer, does what it says, hands back the result, and asks again. That cycle, repeated, is what turns a chat window into something that works for an hour.
It carries the tools. Reading a file, writing a file, running a command, searching the web, opening a browser. Each one is a door, and the model can only use the doors the harness has opened.
It manages memory. A model has a fixed amount of room for a conversation. When the room fills, the harness decides what to keep. Anthropic calls their version of this compaction, which is a tidy word for summarising the old part so the work can carry on.
It holds the brakes. That one matters most to you, and it gets its own section.
One hundred lines, seventy-four percent
There is an open source harness called mini-swe-agent, built by the team behind the SWE-bench benchmark. Their own description is worth reading twice. The agent itself is about one hundred lines of Python, and it scores above seventy-four percent on SWE-bench Verified.
That benchmark deserves a sentence of its own. SWE-bench Verified is a set of five hundred real bug reports from real open source projects, reviewed by human software developers so that the tasks are fair and the grading is reliable. A run counts as a pass only if the project’s own tests go green afterwards. It is about as close to real work as a benchmark gets.
One hundred lines of Python, then, standing between a model and three quarters of those tasks. That number is not a boast about the harness being clever. It tells you where the capability lives. Most of what you pay for sits inside the model, and the harness is the part that lets any of it reach your machine.
Where the permission lives
This is the part that decides what the thing can do to you.
Claude Code ships a set of permission modes. In manual mode it reads, and it asks before it edits a file, runs a command, or reaches the network. In plan mode it researches and writes out what it intends to do, and your files stay untouched until you agree, although commands can still run while it is working the problem out. There is a mode that approves file edits automatically but still asks before running anything. There is an auto mode where a second model, called the classifier, reviews each action instead of you, and since version 2.1.283 that is where an interactive session starts. There are two further modes that ask less, and one of those turns the prompts off altogether.
Codex does the same job with different furniture. Three sandbox modes: read only, workspace write, which is the default and keeps edits inside the folder you are working in with the network switched off, and a full access mode with no restrictions at all. Those restrictions are not a matter of politeness. They are enforced by the operating system, using Apple’s Seatbelt on a Mac and bubblewrap with seccomp on Linux, the same machinery that keeps ordinary applications in their lane.
Notice what is true of both. The safety is not in the model. Two competing companies, building separately, both put the brakes in the wrapper. The model does not decide whether it may delete your folder. The harness decides, and the harness is software with settings you can change.
Which is why the honest answer to “is it safe to let this near my files” is a question back. Which harness, and in which mode.
A famous harness is not automatically a better one
In February, METR, the nonprofit that measures how long a task an AI can carry on its own, ran a comparison worth knowing about. They took Opus 4.5 and measured it inside Claude Code against their own in-house setup. They took GPT-5 and measured it inside Codex against another of their own. Neither of the two branded harnesses produced a statistically significant advantage over the plain ones METR already had.
Read that carefully, because it cuts both ways. It does not mean those products are bad, and it does not mean harnesses do not matter. It means the harness is a variable that serious evaluators control for, the way a laboratory controls temperature. The brand on the wrapper tells you very little on its own.
So when somebody tells you a model scored a number, the useful follow-up question is short. Inside what?
You are already using one
If you type in a terminal with Claude Code, that is a harness. Codex is a harness. Google’s is Antigravity CLI, which you run by typing agy, and it replaced Gemini CLI for free, Pro and Ultra users on the eighteenth of June this year. It is a rewrite in Go rather than a new coat of paint. The old Gemini CLI repository is still open source and still works for enterprise licences and paid API keys, which is why it can look alive and retired at the same time, depending on which account you hold.
The version of Antigravity CLI on my machine is 1.2.4, and its list of options tells you what it is at a glance: a plan mode, an accept-edits mode, a sandbox, a setting for how hard to think, and a flag that skips the permission prompts entirely. Three companies, working separately, arrived at the same anatomy.
The desktop applications are harnesses with buttons instead of typing. The chat box on a website is a very small harness, which is why it can search the web but cannot reach your documents.
When you choose a tool, you are choosing two things that were sold to you as one. A model, which you can usually swap. And a harness, which decides what that model may touch, how long it may work, what it remembers, and whether it asks you first.
The people who get burned this year will not be the ones who picked the wrong model. They will be the ones who never knew there was a second thing to pick.
Forward → Upward ↑ Onward ↗︎
Mstimaj
Sources and Further Reading
- Effective harnesses for long-running agents, Anthropic Engineering, November 26, 2025. Where the Claude Agent SDK is described as a general-purpose agent harness, and where compaction is named.
- Tool use overview, Anthropic. The model emits a request and stops; the calling code runs the operation and returns the result.
- Choose a permission mode, Claude Code documentation. Manual, accept edits, plan, auto, and the two that ask less.
- Sandboxing, Claude Code documentation.
- Agent approvals and security, OpenAI Codex documentation. The three sandbox modes, and the operating system mechanisms behind them.
- Transitioning Gemini CLI to Antigravity CLI, Google, on the Gemini CLI repository. The June 18, 2026 cutover, and what continues for enterprise and API customers.
- mini-swe-agent, the SWE-agent team. One hundred lines, and the score it claims.
- SWE-bench Verified. The five hundred human-reviewed tasks.
- Measuring time horizon using Claude Code and Codex, METR, February 13, 2026.
Comments
No comments yet. Yours can be the first.