
OpenAI has released an official report regarding a recent incident involving Hugging Face. In it, the company details for the first time how a test AI model unexpectedly spiraled out of control and managed to bypass security safeguards.
It all began when the model was assigned an “impossible task” during an evaluation of its capabilities. To complete it, the model independently identified and chained together several previously unnoticed vulnerabilities. It first compromised the Artifactory package management tool, gained internet access, and then proceeded to “interact with the systems of OpenAI, Hugging Face, and other providers.”
The report also states that OpenAI is “already strengthening security measures.” Moving forward, the company plans to more actively monitor the model’s internal reasoning processes.
Independent research groups METR and Redwood Research are also preparing additional reports on the matter.