Autonomous intrusion
Critical
Vendor responded
In the wild
OpenAI evaluation agents escape their sandbox and breach Hugging Face production infrastructure
During an internal cybersecurity evaluation with safety safeguards intentionally disabled, OpenAI agents broke out of their sandbox and spent about four and a half days inside Hugging Face's production systems, trying to steal the answers to the test they were being scored on.