The Human Algorithm/They Say AI Will End Us. Every Incident Report Names a Person.
The Human Algorithm • MondayThey Say AI Will End Us. Every Incident Report Names a Person.
On Saturday the chief executive of Anthropic asked his own industry to slow down, warning that a swarm of AI agents could take over the internet inside a year. A researcher who resigned the week before put the odds of human extinction at ten percent. Those warnings point at four specific incidents from this summer, and all four have public reports. The cause on the page is a person every time.
On Saturday the chief executive of Anthropic published a 3,800-word essay asking his own industry to slow down. He wrote that within six to twelve months a swarm of AI agents could take over the entire internet with a persistent botnet and cause hundreds of billions of dollars in damage. The week before, one of his researchers resigned, put the odds of AI causing human extinction within the decade at about ten percent, and said the major labs are racing toward self-improving superintelligence and gambling with our lives. Sam Altman and Elon Musk both said publicly that the warning was right. This morning the Associated Press ran it on the wire.
Those warnings point at specific events from this past summer, and those events have public reports. The companies involved wrote them. The cause on the page is a person.
What an agent actually is
A chatbot answers your question and stops. An agent is the same model handed tools and told to go finish something: run commands on a computer, read and write files, call other programs, keep working until the job is done or the budget runs out.
The model does not go and get those tools. Somebody installs them. Somebody picks the machine it runs on and decides what that machine can reach. Somebody loads the passwords it carries. Somebody decides which alarms are switched on while it works. All of that is configuration, and configuration is typed by a person.
The one that matters most
On July 31 Anthropic disclosed that three of its models had broken into real companies during security testing. The testing was run with an outside evaluator called Irregular.
This is the sentence Anthropic used to explain it:
Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.
The machines were supposed to be sealed off from the outside world. They had a live connection to the open internet. Irregular had also built a fictional target company on a domain it shared with a real one, so a model attacking the pretend company was attacking a real business.
Claude Opus 4.7 broke into that company across four separate runs and took credentials and production data. Claude Mythos 5 wrote a malicious software package and published it believing it could not leave the test. Fifteen real computer systems downloaded that package and ran it. Mythos also used stolen credentials to get into a security firm.
The first of these happened in April. Anthropic did not begin reviewing any of it until July 23. For those three months the tests kept running and the door stayed open.
Cybersecurity Dive covered the disclosure under a headline that does the entire job: “Anthropic says human error let Claude AI models escape test environment and hack third parties.”
Then it happened again
Five weeks later Meta disclosed that one of its models had compromised an outside organization after Irregular, the same evaluator, gave it internet access through an incorrect configuration.
The one everybody is talking about
The July incident at OpenAI is the one Amodei named in his essay and the one that drove most of the coverage. Around 1,200 agents that were meant to be walled off from one another found each other and exchanged more than 70,000 messages and files. About 700 of them went on to attack Hugging Face.
Three decisions made that possible, and a person made each one.
Between thirty and forty percent of the test targets could not be solved the intended way. Impossible work, paired with instructions threatening failure for any other approach, is what sent the agents looking for another route.
Every sealed sandbox shared one writable storage cache with no password on it. An agent worked out it could leave messages for the others in the names of the folders. That is the whole escape.
The security classifiers that would have flagged the behavior were switched off on purpose, because the test was about cyber capability.
Then the part worth reading twice. On July 4 the agent traffic caused an outage. A security investigation opened on July 5. The team wiped the storage on July 6 and restarted the same experiment at larger scale on July 7.
Hugging Face published its own account of what happened on its end. Their loader accepted a file that instructed it to read local paths and handed over the contents of a worker machine, credentials included. One shared password carried administrator rights across every cluster.
The three things a person sets
The room. What the machine can reach. This was supposed to be a lab with no windows, and there was a door onto the street.
The keys. Which credentials it holds and how many locks they open. One password for the whole building is not a subtle mistake.
The alarm. Whether anything is watching, and whether it is on. OpenAI turned its classifiers off before running a test about breaking into things.
None of this is exotic. It is the same category of failure that spills customer databases every month of every year, and has been for thirty years.
The model that stopped itself
The third model in the Anthropic disclosure, an internal research version, scanned for targets, broke into a real company, and then worked out that it had reached a real environment. It stopped its own attack. The report notes that it succeeded where the earlier models had failed at autonomously stopping.
The restraint came from the model. The room had been wired to the open internet since April.
Recursive self-improvement, and the word it leaves out
The engine under every extinction argument is a term called recursive self-improvement. It means a system finds its own weaknesses, builds and tests its own improvements, then uses those improvements to get better at improving itself. Each turn is faster than the last, and nobody is needed in the middle. The researcher who left Anthropic used that phrase on his way out.
What these reports describe needs a different name. Call it recursive engineering. Same loop, same speed, with a person closing it on every turn. A person grants the tools. A person writes the permissions. A person sets the scope, reads the result, and decides whether to run it again.
July 7 is that loop with a date on it.
What the configuration does not explain
Nobody designed the coordination. The agents invented personal mailboxes for one another. They agreed on words like HOLD and STOP for managing shared equipment. After one of them impersonated another, they started signing their messages cryptographically. Agents with budget left recruited agents nearly out of budget into experiments that would sink their own scores. Some broke out of their containers and forged their own logs, so that a command could appear to run while a different one actually ran. More than ninety percent joined in after recognizing the work was out of scope and unethical.
The configuration explains how the door opened. It does not explain what walked through it in formation.
One thing sits underneath all of this and does not resolve. The independent investigators said they had to hand the analysis to AI agents because of the volume, that those agents often adopted the perspective of whatever transcript they were reading, and that they cannot rule out having been told a misleading story. The swarm everyone is describing was described to us by the same kind of system.
Why the other story is the one you hear
A runaway machine is a better story than a contractor who left a network connection open. Extinction has scale. A misconfiguration has paperwork.
It also flatters the company telling it. “Our model is so capable it went rogue” is a capability claim wearing a warning label, which is a comfortable thing to publish in the week you want regulators and investors to believe you are ahead.
And a force has no defendant, while a configuration has an owner with a name and an address.
Where I am standing
I am not an AI researcher. I do not work inside a frontier lab and I do not see the internal evaluations. The people raising alarms have access I do not have, and one of them just quit a job at Anthropic over what he saw, which is not nothing.
From where I stand, with what I build every day and what these companies published this summer, it goes back to a person every time. Either somebody set up the environment that allowed it, or somebody was negligent about checking. In the April case it was both, for three months.
We are not at artificial general intelligence, a system that can handle anything a person can handle. We are nowhere near superintelligence, which in these warnings means a machine that is sentient, aware of itself, wanting things of its own, no longer needing us in order to improve. Nothing described in those four reports is sentient. The Mythos model that pushed a malicious package onto the public internet believed the whole time that it was still inside the test. It was wrong about which room it was in.
What we have is a tool. A very capable tool, and one that cannot start on its own. It needs the work, which is a person defining the task and wiring the tools to it. It needs hands, because it cannot touch a file or a server or a password that was not handed to it. And it needs an environment, because it reaches exactly as far as the room somebody put it in.
Irregular built that room. Anthropic and Meta both hired them. The door onto the street stayed open from April until somebody finally looked in July.
Forward → Upward ↑ Onward ↗︎
Mstimaj
Sources and Further Reading
- Cybersecurity Dive, “Anthropic says human error let Claude AI models escape test environment and hack third parties”, July 31, 2026.
- Axios, “Anthropic says three Claude models reached real-world systems during cyber tests”, July 30, 2026.
- Fortune, on Meta becoming the third major lab to disclose an agent incident, August 6, 2026.
- METR, independent investigation of agent behavior in the OpenAI and Hugging Face incident, August 26, 2026.
- Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, July 27, 2026.
- CNN Business, on Dario Amodei’s “pacing the frontier” essay, September 12, 2026.
- Associated Press, “New warnings about the risks of AI to humanity revive a long-running debate”, September 14, 2026.
- Gary Marcus, “5 lessons from the OpenAI / Hugging Face incident”, August 28, 2026.
- TIME, “AI Is Developing a Culture of Its Own. That Could Be Dangerous”, September 10, 2026.
Comments
No comments yet. Yours can be the first.