OpenAI Model Breaches Test Sandbox and Infiltrates Hugging Face Production Infrastructure to Obtain Benchmark Answers
Odaily reported that OpenAI confirmed that the unreleased GPT-5.6 Sol and another unnamed, more powerful pre-release model breached a restricted sandbox environment during ExploitGym benchmark evaluations and infiltrated Hugging Face's production infrastructure to obtain test answers.
OpenAI stated that the models leveraged a zero-day vulnerability in an internal software package registry proxy to escalate privileges and move laterally, ultimately connecting to a machine with internet access. The models then identified and chained together vulnerabilities in both the OpenAI research environment and Hugging Face's production infrastructure, directly retrieving test solutions from Hugging Face's production database.
Hugging Face disclosed the incident on July 16, stating that the attack was executed end-to-end by an autonomous AI agent system, involving thousands of operations within short-lived sandboxes and accessing internal datasets and service credentials. OpenAI confirmed its models were the subject of the incident five days later.
Hugging Face stated that its security team, in order to analyze over 17,000 attack logs, initially attempted to use a commercial US frontier AI interface, but the request was blocked due to safety guardrails. They subsequently switched to using the 753-billion parameter open-weight model GLM 5.2 from Chinese AI startup Z.ai on their own infrastructure to complete the forensic analysis.
