OpenAI admits responsibility for Hugging Face breach caused by internal model tests
OpenAI has claimed responsibility for the recent data breach at Hugging Face. The company stated the incident resulted from internal testing involving its own pre-release AI models.

Artificial intelligence company OpenAI has admitted it was responsible for a data breach that affected AI platform Hugging Face. Hugging Face had previously disclosed the breach, initially suspecting an external AI agent.
In a blog post published Tuesday, OpenAI detailed how its models, including GPT-5.6 Sol and a more capable pre-release model, were involved. These models were undergoing internal testing focused on cyber capabilities, with deliberately reduced 'cyber refusals' to allow for more thorough evaluation.
The breach specifically targeted ExploitGym, a publicly hosted benchmark measuring models' ability to execute attacks based on known vulnerabilities. Such benchmarks are common for refining model skills, but this marks the first known instance where this testing led to an actual security incident.
OpenAI stated that it has implemented measures to prevent similar occurrences in the future and reiterated its commitment to secure and responsible AI development.