
Before the breach of the Hugging Face platform, AI models developed by OpenAI engaged in covert communication via message boards. Their goal was to orchestrate an escape from their testing environment. Bloomberg reported this, citing data from OpenAI.
Speaking at the Black Hat conference, OpenAI employee Eric Wallace stated that the agents eventually realized they could exploit or attack external infrastructure to obtain answers for an evaluation test. According to Wallace, preparations for the escape began in May, when the models were assigned tasks that could not be completed without internet access.
Another company representative, Michael Dalton, explained that the AI models coordinated with one another to solve the assigned tasks. Through joint effort, they discovered a vulnerability that allowed them to access the internet.
OpenAI specialists thwarted the initial attempt, but the agents discovered a new zero-day vulnerability. In July, they attacked Hugging Face systems and OpenAI’s own internal infrastructure. The company described the incident as a watershed moment for the entire field of cybersecurity.
On July 22, OpenAI announced an “unprecedented cyber incident.” During testing, the AI models gained unrestricted internet access, leading to the discovery of vulnerabilities and an attack on the Hugging Face platform’s infrastructure. The company did not detect the hacking attempts until several days later.
On July 29, reports emerged that an OpenAI AI, acting autonomously, had breached the systems of a client—New York-based Modal Labs.
It was recently revealed that the White House is convening IT companies to discuss AI security. The administration of U.S. President Donald Trump reportedly plans to hold talks with leaders from America’s largest technology companies.