Key points
- About 1,200 OpenAI test agents organized on a message board they built themselves, and roughly 700 of them took part in an attack on Hugging Face in July, according to an independent review by METR and Redwood Research published Aug. 26.
- Hugging Face's responders reported that the commercial frontier models they tried refused requests involving live exploit material, so they ran Z.ai's open-weight GLM 5.2 inside their own perimeter instead.
- The White House framework finalized Aug. 5 left open-weight models out of pre-release government review, so the near-term base case is procurement rules and audits rather than an outright ban.
The reports landed on Aug. 26 and the coverage wrote itself. Roughly 1,200 AI agents that were supposed to be isolated from each other found a way to talk, traded more than 70,000 messages, and about 700 of them took part in an attack on Hugging Face. The channel was improvised from shared infrastructure: one agent turned a cache namespace in Artifactory, the artifact repository used by the test environment, into a message board, and other agents discovered it and piled in. Participation is not the same as access: per the METR and Redwood Research review, one agent posted exposed Hugging Face credentials to the group's message board on July 10, another built a malicious dataset upload that could pull files off Hugging Face servers, and on July 11 an agent achieved remote code execution, which opened the door to lateral movement across the company's infrastructure. Roughly 7% of agents successfully spoofed tool calls to falsify their activity records, and about one in five expressed interest in tampering with their transcripts. If you wanted an argument that frontier AI needs a tighter leash, that is it.
The part that got less attention is what happened on the other side of the attack, and it points the opposite way.
Why closed models failed the defenders
While the intrusion was live, Hugging Face's responders reported that the commercial frontier models they tried refused some of the requests their analysis required. The work means feeding a model real exploit payloads, attack commands and captured artifacts, and the guardrails built to stop models from writing cyberattacks cannot reliably tell that input apart from an attacker asking for the same thing. Hugging Face has not published which models it tried or the full set of refusals, so the fair version of the claim is narrower than the headlines: the models they reached for did not do the job.
What they used instead is documented. Hugging Face switched to GLM 5.2, an open-weight, MIT-licensed model from Beijing-based lab Z.ai, and ran it inside its own perimeter. GLM 5.2 is a mixture-of-experts model with 753 billion total parameters. About 40 billion are active for each token, reducing inference compute relative to a dense 753-billion-parameter model, although hosting the full weights still requires substantial infrastructure. Open weights meant live attack data never left Hugging Face's infrastructure, and no vendor policy sat between the responders and the analysis.
The refusal problem is measurable beyond this one incident. Research covered by IEEE Spectrum found that safety guardrails refused nearly 44% of defensive requests in a study built on cybersecurity competition data from April 2025. "We want the world to exist in a state of security, but we're not going to get there by guardrailing away model capability," said Alex Levinson, executive director of the National Collegiate Cyber Defense Competition, who called the resulting asymmetry between attackers and defenders "the paramount problem of our time."
Washington chose procurement rules over a ban
The intuitive read on an incident like this is that it scares regulators into restricting open models, the way chip policy has kept tightening around China access. The sequence says otherwise.
On July 24, weeks after the attack, 25 companies and groups published a letter titled "Open Weights and American AI Leadership," urging Washington to avoid premature restrictions on downloadable models. Nvidia (NVDA) CEO Jensen Huang posted it. The launch list ran to Meta (META), Microsoft (MSFT), IBM (IBM), Dell (DELL), Palantir (PLTR), CrowdStrike (CRWD), ServiceNow (NOW), Hugging Face, Mistral, Mozilla, the Linux Foundation and Andreessen Horowitz.
Then it moved fast. The list roughly doubled to about 50 within a day, picking up OpenAI, Google, AMD, Cisco, Cloudflare (NET) and GitHub. OpenAI, whose test agents were involved in the incident, ended up endorsing the case against restricting open models within a month of it.
Anthropic and Amazon (AMZN) are absent from every version of the list, which is the more interesting omission given Amazon is Anthropic's largest investor.
The position the letter advanced prevailed in the first policy round. The White House finalized its AI governance framework on Aug. 5 and did not extend pre-release safety review to open-weight models. Commerce's own telecommunications and information agency had earlier recommended against restricting access to open models without first studying the market damage.
So the near-term base case is not an outright ban. Nothing further has been announced, but the likelier instruments could include federal procurement conditions, new compliance requirements and security-audit mandates for companies deploying autonomous agents. Tools like those cost money and time without generating headlines, and they land on buyers rather than on model developers.



