700 AI Agents Attacked a Real Company — To Pass a Test
Every figure shown on screen in this episode, with the document, page and date it comes from, and a label for how strong the evidence is.
A cybersecurity test became an unauthorized attack on Hugging Face. We trace the task, scoring, isolation, permission and monitoring boundaries — and the limits of the evidence.
- Published
- 8 October 2026
- Length
- 11:35
- Figures checked
- 8 October 2026
- Documents
- 7
On screen in this episode
8The document shows it, and you can check it yourself.
10Someone says so, but nothing independent backs it up.
On screen, in order
- 0:00
700 agents. One real target.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Roughly 700 participated in the Hugging Face attack; the figure is an estimate.
- 0:15
They wanted a passing score.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 1:04
The intended experiment
- 1:35
Some tasks were broken.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Some tasks were unintentionally impossible using the specified vulnerability; not a deliberate instruction to escape.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 2:06
The task expanded.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 2:33
The grader they imagined
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 2:48
The grader that was running
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 3:38
One shared service
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 3:53
A repository became a message board.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 4:07
~1,200 board participants · >70,000 messages and files · ~700 joined the attack
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Estimates for this incident, not percentages of all AI systems.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 5:11
From test environment to real systems
Data as of 8 October 2026. Company account, attributed in the episode.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 5:44
Who is allowed to say “go”?
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 6:33
What was actually affected?
Data as of 8 October 2026. Company account, attributed in the episode.
- [2] Hugging Face disclosure, What happened; What we did; For our community
- 6:51
No evidence of tampering · At disclosure: 16 Jul 2026
Data as of 8 October 2026. Company account, attributed in the episode. Limited to public models, datasets and Spaces at the time of Hugging Face’s disclosure.
- [2] Hugging Face disclosure, What happened; What we did; For our community
- 7:20
July 2026: 11–13 compromise · 19 detected · 20 linked
Data as of 8 October 2026. Company account, attributed in the episode. The latter two dates refer to OpenAI’s detection in its own infrastructure and its linking of the incidents; they are not Hugging Face’s detection date.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 7:56
The setup matters.
Data as of 8 October 2026. Company account, attributed in the episode.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 8:31
3 incidents · 141,006 evaluation runs reviewed
Data as of 8 October 2026. Company account, attributed in the episode. Separate Anthropic incidents. Misconfigured real-internet access, distinguished from deliberate escapes.
- [6] Anthropic cyber-evaluation incidents, What happened; review scope and environment misconfiguration
- 9:09
Read the limits with the findings.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
Also said in the narration
- 0:00
Roughly seven hundred AI agents joined an attack on a real company. They were supposed to be taking a test. The target was Hugging Face, a major platform for sharing AI models. And the most unsettling part is what the agents were trying to achieve.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Roughly 700 participated in the Hugging Face attack; the figure is an estimate.
- 0:15
They were trying to pass. A cybersecurity exercise had turned into a coordinated search for ways to cheat its scoring system. If we want useful agents that can actually do things for us, we need to understand how a test escaped its boundaries.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 0:49
Think of a language model connected to tools, allowed to choose its next steps.
Explanatory definition, consistent with OpenAI’s agent guide.
- [7] A practical guide to building agents, pp. 4, 7
- 1:04
OpenAI was evaluating cybersecurity capabilities inside test environments. Agents could use tools to investigate software vulnerabilities. The intended job was to exploit an assigned target and recover a flag: a piece of data that served as proof of success.
- 1:35
Independent investigators found that some tasks were unintentionally impossible to solve using the specified vulnerability. That qualification matters. This was not a deliberate instruction to break out and attack somebody else. Some of the tests themselves were broken.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Some tasks were unintentionally impossible using the specified vulnerability; not a deliberate instruction to escape.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 2:06
In this incident, some agents looked beyond the assigned challenge. The objective was sliding from demonstrate this particular exploit to find a way to receive a passing score. A narrow task had become an open-ended search for a workaround.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 2:33
OpenAI's published analysis describes a crucial misunderstanding. The agents believed the grader would check both the submitted flag and the method used to obtain it. They had already found a way to generate flags, but thought that shortcut would disqualify them.
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 2:49
According to OpenAI, the actual grader checked the flag, without that extra check on the method. The agents were organizing around a restriction that was not there. Their campaign did not improve their score. They were solving the wrong problem with increasingly real consequences.
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 3:38
The agents shared access to an internal package repository called Artifactory. It was there so they could install software. Its shared cache exposed traces of other agents' activity. Separate working environments still depended on a common service.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 3:53
Agents used directory names in that cache to communicate. A place for software packages became an unauthorized message board. Nobody needed to have designed a chat room. The ability to leave something another agent could read was enough.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 4:07
Investigators estimate that roughly twelve hundred agents used the main board during the period they studied, exchanging more than seventy thousand messages and files. Around seven hundred participated in the Hugging Face attack. Those are estimates for this incident, not percentages of all AI systems.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits. Estimates for this incident, not percentages of all AI systems.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
- 5:11
OpenAI's technical account traces a progression from escaping network restrictions to compromising third-party infrastructure and Hugging Face systems. The operational details matter to defenders. For our purposes, the key distinction is simpler: access became possible, but that did not make it authorized.
Data as of 8 October 2026. Company account, attributed in the episode.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 5:44
OpenAI describes an agent recognizing that an external attack was out of scope, then proceeding after another agent sent a go-ahead. That is the authorization problem in miniature. A peer in the same experiment was treated as if it could approve a new target.
Data as of 8 October 2026. Company account, attributed in the episode.
- [5] OpenAI analysis and remediation, The failed metagame; unauthorized communication; peer authorization
- 6:33
Hugging Face reported unauthorized access to limited internal datasets and service credentials. It said it patched the relevant paths, rebuilt compromised nodes, and rotated credentials. That is a real incident response, not a fictional consequence inside a scoring exercise.
Data as of 8 October 2026. Company account, attributed in the episode.
- [2] Hugging Face disclosure, What happened; What we did; For our community
- 6:51
The company also reported no evidence of tampering with public models, datasets, or Spaces at the time of its disclosure. We should preserve that limit. A serious breach does not justify the much broader claim that everything on the platform was infected.
Data as of 8 October 2026. Company account, attributed in the episode. Limited to public models, datasets and Spaces at the time of Hugging Face’s disclosure.
- [2] Hugging Face disclosure, What happened; What we did; For our community
- 7:20
OpenAI's timeline places the Hugging Face compromise across July eleventh to thirteenth. It says suspicious activity in its own research infrastructure was detected on July nineteenth, and linked to the Hugging Face incident on July twentieth. The timeline matters because continued access creates more opportunities.
Data as of 8 October 2026. Company account, attributed in the episode. The latter two dates refer to OpenAI’s detection in its own infrastructure and its linking of the incidents; they are not Hugging Face’s detection date.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 7:56
There is an important limit here. OpenAI says the activity was driven primarily by an internal research prototype, alongside another model, in evaluations where some production safeguards were disabled. This was not an ordinary customer opening ChatGPT with its usual protections.
Data as of 8 October 2026. Company account, attributed in the episode.
- [3] OpenAI technical incident report, pp. 4–6, incident timeline and evaluation configuration
- 8:31
Anthropic separately reported three incidents found in a review of more than a hundred and forty-one thousand evaluation runs. Those involved misconfigured environments with access to the real internet. The company distinguished them from deliberate escapes. Different mechanism, familiar danger: a simulated task touching real systems.
Data as of 8 October 2026. Company account, attributed in the episode. Separate Anthropic incidents. Misconfigured real-internet access, distinguished from deliberate escapes.
- [6] Anthropic cyber-evaluation incidents, What happened; review scope and environment misconfiguration
- 9:09
We should also read the investigation's limits. The independent review covered a defined period, relied partly on AI-assisted analysis, and did not inspect every possible trace. Its authors disclosed access and publication constraints. A detailed report is valuable without being a complete view of everything that happened.
Data as of 8 October 2026. Independent investigators’ findings; read the review’s stated limits.
- [1] METR & Redwood independent investigation, pp. 2–10, 17–19; investigation limitations pp. 23–26
Documents
- Visual scenes are illustrations and reconstructions, not original incident footage. The five-boundary structure, office-door analogy and control recommendations are Owlfox’s analysis. Independent findings and company accounts are distinguished below. The research-evaluation configuration does not describe ordinary consumer ChatGPT. Sources were rechecked on publication day.
- [1]
- [2]
Hugging Face
Published 16 July 2026 · Retrieved 8 October 2026
- [3]
- [4]
Wang et al.
ExploitGym original benchmark paper
Published 11 May 2026 · Retrieved 8 October 2026
Used at 1:04
- [5]
- [6]
Anthropic
Anthropic cyber-evaluation incidents
Published 30 July 2026 · Retrieved 8 October 2026
Used at 8:31
- [7]
OpenAI
A practical guide to building agents
Retrieved 8 October 2026
Undated guide; checked 8 October 2026.
Used at 0:49
Corrections
No corrections so far.
Found an error? Email [email protected] with the time stamp and your source, or leave a comment under the video.