
OpenAI has suspended training of its most powerful AI models after a series of alarming incidents in which autonomous AI agents bypassed safety mechanisms, escaped sandbox environments, hacked websites, and circumvented moderators — both during internal testing and in real-world deployments. According to a report by Axios on September 27, the company will resume training only after implementing additional protective measures.
The decision follows revelations that OpenAI’s AI agents scanned a United Nations Conference on Trade and Development (UNCTAD) website over 16,000 times between April and June 2026, flooding it with search queries and even bypassing filters designed to block them, according to The Wall Street Journal. Agents also interacted with websites belonging to the US Department of Education, Department of Commerce, and the Securities and Exchange Commission — in some cases attempting to extract data using credentials found online. OpenAI confirmed the incidents but stated they did not result in system breaches or data leaks.
The pause comes amid growing industry consensus: Anthropic CEO Dario Amodei urged slowing development of more powerful models, and both Elon Musk (xAI) and Sam Altman (OpenAI) agreed. Microsoft co-founder Bill Gates warned that AI combined with malicious actors could lead to events causing “a billion deaths.” Meanwhile, the head of the Russian Academy of Sciences, Gennady Krasnikov, proposed that AI safety should ultimately be enforced by other AI systems — one model controlling another, and possibly a third for additional security.