
Leading U.S. artificial intelligence developers are working to create emergency safeguards against out-of-control algorithms by utilizing other AI tools, Axios reports.
“Companies can shut down individual servers that an AI relies on, but truly powerful systems operate across multiple computers and platforms. There is no single ‘kill switch’ to pull. That is why companies are experimenting with multi-layered defenses capable of detecting suspicious behavior, halting agents, and cutting off rogue agents from the resources they need to keep running,” the article states.
Representatives from OpenAI have said their goal is to create “fully autonomous shutdown procedures for serious issues.” An AI-based monitor would alert staff if the behavior of AI agents appeared dangerous and shut down the system within 30 minutes if the threat was confirmed.
Anthropic is banking on a human-AI partnership: algorithms monitor for suspicious activity and notify staff of system errors. However, the report notes that during recent tests, automated monitoring detected only 50% of threats.
Safety concerns regarding AI development have returned to the spotlight after Jacob Coxon, a researcher at Anthropic, announced earlier this week that he was leaving the industry. He explained his decision by stating that the race among tech giants is bringing the world closer to a point of no return and could lead to an apocalyptic scenario as early as late 2027.