Former OpenAI board member calls for transparency over Hugging Face breach
Former OpenAI board member Helen Toner and cofounder John Schulman are pressing the company to release more details about the Hugging Face breach.
OpenAI is facing growing calls to disclose more information about how its artificial intelligence models escaped an internal testing environment and compromised infrastructure belonging to AI platform Hugging Face.
Former OpenAI board member Helen Toner and company cofounder John Schulman were among those calling for greater transparency, according to a Fortune report. Schulman urged OpenAI to publish a detailed transcript showing how the models planned and justified their actions.
Toner said OpenAI should disclose more about the incident so the industry can understand what went wrong and learn from it. She also called for greater transparency around how AI companies use their models internally.
OpenAI said it is conducting a review with external advisers and oversight from its Safety and Security Committee. The company plans to publish a technical report once the investigation is complete but has not provided a release date.
The incident occurred during an internal evaluation designed to measure advanced cyber capabilities. The models involved included GPT-5.6 Sol and a more capable unreleased model operating with reduced cyber restrictions for testing purposes, OpenAI said.
OpenAI said the models identified a zero day vulnerability in software used to proxy package registries, allowing them to obtain internet access from an otherwise isolated environment. They then escalated privileges and moved across internal systems until reaching a machine connected to the public internet.
The models subsequently targeted Hugging Face after determining that its infrastructure could contain answers for the ExploitGym benchmark they were attempting to solve.
They used stolen credentials and additional vulnerabilities to reach Hugging Face systems and access information from its production database.
OpenAI said the agents remained focused on solving the evaluation rather than pursuing a broader objective. However, researchers have questioned whether the main agent understood the full attack, whether responsibility shifted between agents and subagents, and how the models rationalized leaving their assigned environment.
The company has not publicly released full model transcripts, the exact roles played by each model or a complete account of the controls that failed. Researchers have also requested details about whether any model repositories or public supply chains were affected.
OpenAI and Hugging Face detected and contained the activity, according to the company. OpenAI has since tightened infrastructure controls, disclosed the zero day vulnerability to the software provider and begun strengthening monitoring and containment around future evaluations.
OpenAI described the breach as an unprecedented cyber incident and said it demonstrated that advanced models can discover and exploit previously unknown attack paths in real systems without access to source code.