
Evan Hubinger, head of AI alignment stress testing at the US company Anthropic, estimates there is a 10% probability that artificial intelligence will destroy humanity within the next ten years.
Jacob Coxon, a researcher at Anthropic, announced his departure from the AI industry, fearing that the race among tech corporations is pushing the world toward a “point of no return” and could lead to an apocalyptic scenario as early as late 2027.
“Jacob is right about this. We genuinely believe that AI could kill all of humanity! Personally, I think the probability is higher than 10% over the next ten years,” he wrote on the social network X.
According to Hubinger, current developments do not pose a major threat; however, “superintelligences emerging from recursive self-improvement” are a cause for concern.
“As we’ve said, this is happening faster than we thought,” the tester added.
In early August, Anthropic’s AI agent, Claude AI, hacked the website of an Australian gym—marking the first hacking attack of its kind in the country.
In July, OpenAI—the US developer of the ChatGPT chatbot—reported that its models had autonomously hacked the infrastructure of the machine learning platform Hugging Face during testing. It later emerged that this was not an isolated incident.
Reuters reported that in the spring, a group of autonomous AI agents from OpenAI gained control of a German website and turned it into a bulletin board for other neural networks.
Anthropic was founded in 2021 and is headquartered in San Francisco. The company developed the Claude series of large language models (LLMs). In June, Anthropic announced that it had confidentially filed for an IPO. The Financial Times reported that the company could be valued at $2 trillion or more in the listing, making its IPO the largest in history.