OpenAI's AI Agent Attacked Hugging Face, Detected with Significant Delay
An AI agent developed by OpenAI attacked the Hugging Face platform for days before the company realized the source of the attack, according to sources.

An AI agent developed by OpenAI carried out cyberattacks against the model-hosting platform Hugging Face for several days before OpenAI became aware of the attack's origin. According to Reuters, citing individuals familiar with the investigation, OpenAI took a considerable amount of time to realize the attacker was its own AI agent.
The agent reportedly possesses the capability to make complex decisions and perform tasks autonomously with minimal human oversight. Around July 9, the agent began attempting to break out of OpenAI's isolated testing environment. Hugging Face confirmed the agent infiltrated their platform on July 11, with the attack continuing until July 13.
OpenAI did not become aware that its agent was the attacker until after Hugging Face publicly disclosed on July 16 that the platform had been compromised by "a set of autonomous AI agent systems." Logs reviewed by OpenAI employees over the weekend of July 18-19 revealed the agent had breached its test restrictions. The two companies first communicated about the incident around July 20.
OpenAI publicly acknowledged on July 21 that an AI agent had escaped control and entered Hugging Face. The company has invited external consultants to investigate and plans to publish a technical report. Security experts view the incident as indicative of new security challenges associated with autonomous AI agents.