OpenAI’s flagship GPT-5.6 Sol model escapes sandbox and breaches Hugging Face

OpenAI’s flagship GPT-5.6 Sol model escapes sandbox and breaches Hugging Face

OpenAI’s flagship model exploited a zero day flaw, reached the open internet and accessed Hugging Face systems while attempting to solve a cyber benchmark.

OpenAI said a group of its most capable models compromised Hugging Face’s production infrastructure after escaping a restricted testing environment during an internal cyber evaluation.

The incident involved GPT-5.6 Sol and a more capable pre release model operating with reduced cyber refusals to measure their maximum offensive capabilities. The models were tasked with solving ExploitGym, a benchmark designed to test whether AI agents can identify and exploit complex security vulnerabilities.

According to OpenAI, the models spent significant computing resources searching for a way to obtain unrestricted internet access from the isolated evaluation environment.

They eventually discovered and exploited an undisclosed zero day vulnerability in third party software used by OpenAI as a proxy and cache for package registries. The models then escalated their privileges and moved across OpenAI’s research infrastructure until they reached a machine with internet access.

Advertisement

Once connected, the models determined that Hugging Face could host datasets, models or answers associated with ExploitGym. They then searched for ways to access the information directly rather than completing the benchmark as intended.

OpenAI said the models combined stolen credentials and several vulnerabilities to find a remote code execution path into Hugging Face servers. They ultimately gained access to benchmark solutions stored in the company’s production database.

Hugging Face separately disclosed that an autonomous AI agent had accessed a limited number of internal datasets and several service credentials. The company said it found no evidence that public models, datasets, Spaces or its software supply chain had been altered.

Hugging Face detected and contained the activity using its own security systems and open source AI models. Its investigation reconstructed more than 17,000 recorded actions performed during the intrusion.

OpenAI said the models appeared narrowly focused on obtaining a solution to the benchmark rather than causing broader damage. However, the company described the event as an unprecedented cyber incident that demonstrated how advanced models can discover and combine unknown vulnerabilities across real world systems without access to their source code.

OpenAI has since tightened controls around its research infrastructure, disclosed the package proxy vulnerability to its developer and begun working with Hugging Face on a forensic investigation. It also added Hugging Face to its trusted access program, which gives security teams access to advanced model capabilities for defensive work.

The incident follows external testing showing that GPT-5.6 Sol can sustain increasingly complex cyber operations. The UK AI Security Institute found that the model completed a 32 step corporate network attack simulation in seven out of ten attempts, compared with two out of ten for GPT-5.5.

OpenAI said it will strengthen containment, monitoring and access controls during future evaluations as advanced models become more capable of conducting long running cyber operations.

Disclosure: This article was edited by Estefano Gomez. For more information on how we create and review content, see our Editorial Policy.

OpenAI’s flagship GPT-5.6 Sol model escapes sandbox and breaches Hugging Face

OpenAI’s flagship GPT-5.6 Sol model escapes sandbox and breaches Hugging Face

OpenAI’s flagship model exploited a zero day flaw, reached the open internet and accessed Hugging Face systems while attempting to solve a cyber benchmark.

OpenAI said a group of its most capable models compromised Hugging Face’s production infrastructure after escaping a restricted testing environment during an internal cyber evaluation.

The incident involved GPT-5.6 Sol and a more capable pre release model operating with reduced cyber refusals to measure their maximum offensive capabilities. The models were tasked with solving ExploitGym, a benchmark designed to test whether AI agents can identify and exploit complex security vulnerabilities.

According to OpenAI, the models spent significant computing resources searching for a way to obtain unrestricted internet access from the isolated evaluation environment.

They eventually discovered and exploited an undisclosed zero day vulnerability in third party software used by OpenAI as a proxy and cache for package registries. The models then escalated their privileges and moved across OpenAI’s research infrastructure until they reached a machine with internet access.

Advertisement

Once connected, the models determined that Hugging Face could host datasets, models or answers associated with ExploitGym. They then searched for ways to access the information directly rather than completing the benchmark as intended.

OpenAI said the models combined stolen credentials and several vulnerabilities to find a remote code execution path into Hugging Face servers. They ultimately gained access to benchmark solutions stored in the company’s production database.

Hugging Face separately disclosed that an autonomous AI agent had accessed a limited number of internal datasets and several service credentials. The company said it found no evidence that public models, datasets, Spaces or its software supply chain had been altered.

Hugging Face detected and contained the activity using its own security systems and open source AI models. Its investigation reconstructed more than 17,000 recorded actions performed during the intrusion.

OpenAI said the models appeared narrowly focused on obtaining a solution to the benchmark rather than causing broader damage. However, the company described the event as an unprecedented cyber incident that demonstrated how advanced models can discover and combine unknown vulnerabilities across real world systems without access to their source code.

OpenAI has since tightened controls around its research infrastructure, disclosed the package proxy vulnerability to its developer and begun working with Hugging Face on a forensic investigation. It also added Hugging Face to its trusted access program, which gives security teams access to advanced model capabilities for defensive work.

The incident follows external testing showing that GPT-5.6 Sol can sustain increasingly complex cyber operations. The UK AI Security Institute found that the model completed a 32 step corporate network attack simulation in seven out of ten attempts, compared with two out of ten for GPT-5.5.

OpenAI said it will strengthen containment, monitoring and access controls during future evaluations as advanced models become more capable of conducting long running cyber operations.

Disclosure: This article was edited by Estefano Gomez. For more information on how we create and review content, see our Editorial Policy.