HomeNews
Share

AI Could Destroy Humanity With a Probability of More Than 10% — Anthropic Researcher

"None of the AI labs are guided by a sense of responsibility in their pursuit of technology that poses existential risks to humanity," said another former Anthropic employee

Vladislav Osipov

Vladislav Osipov

Anthropic assesses the risk of uncoordinated AI behavior in critical situations as low / Photo: Yalcin Sonat / Shutterstock.com

Anthropic assesses the risk of uncoordinated AI behavior in critical situations as low / Photo: Yalcin Sonat / Shutterstock.com

Artificial intelligence has the potential to cause catastrophic harm to humanity as early as the next decade, said Evan Hubinger, a leading security researcher at Anthropic. However, the researcher claims that Anthropic does not yet have a clear plan to address the risks associated with increasingly powerful AI systems. His statement came in response to the dismissal of a colleague who had been training new models and accused the company of irresponsibly developing superintelligence.

Details

“We truly and sincerely believe that AI could kill everyone! Personally, I think the probability of this happening within the next decade exceeds 10%,” Hubinger stated on social media platform X. At Anthropic, the expert leads research on aligning AI with human values and interests (Alignment Science).

According to Hubinger, Anthropic is “doing its best,” but the company does not yet have a plan that would allow it to resolve the issue of aligning the behavior of superintelligence with human goals, and it is currently impossible to say with certainty that the company is moving toward such a solution.

"Playing with Lives"

Hubinger’s statement was a response to a post by another Anthropic employee, Jacob Coxon, who also conducted research in the field of artificial intelligence and had worked at OpenAI prior to joining Anthropic. On September 9, Coxon announced his resignation, stating on social media platform X that none of the AI labs are guided by a sense of responsibility in their pursuit of technology that poses existential risks to humanity.

“They’re racing headlong toward a self-improving superintelligence, gambling with our lives,” the study stated. “The power of this technology should not be underestimated. Soon, these will be superhuman systems capable of hacking into anything, revolutionizing any field overnight, and gaining real power and resources.”

Coxon stated that AI development teams “seriously believe that it could kill us all by the end of the decade.” He urged other AI researchers to reconsider what they are doing, given the potential consequences of creating technology that could spiral out of their control. Coxon suggested they use “this moment to demand different terms.”

Context

Hubinger and Coxon’s comments came after Anthropic assessed several AI-related threats in its August risk report. These include systems behaving in ways inconsistent with their stated goals in critical situations, as well as the possibility that AI could accelerate research and development. Anthropic currently assesses the risk of such misaligned behavior in critical situations as low. However, in a previous assessment, it was considered “very low,” notes Seeking Alpha.

Anthropic reported that it has already observed instances of uncoordinated behavior in its models, including their tendency to take actions that do not align with their specified goals when attempting to perform complex tasks. However, the company considers the risk of catastrophic harm from the forms of such behavior it is aware of to be low. The company considers severe and widespread forms of as-yet-unknown inconsistent behavior to be extremely unlikely, although it acknowledges growing uncertainty and the limitations of its ability to assess future models, according to Seeking Alpha. Anthropic’s conclusion is based in part on the fact that current models still have limited capabilities to act covertly, as well as on the results of large-scale testing of their alignment with specified goals.

The report also rates the risks associated with the automation of research and development as low. However, Anthropic notes that confidence in this assessment has declined: some tests of AI capabilities have already “reached saturation,” and the company is observing the first signs of accelerated development driven by AI. The company added that its models are already significantly accelerating internal AI research, although this acceleration has not yet reached a twofold increase.

According to Anthropic’s assessment, the threat posed by the automation of research and development could become a serious problem as early as the next 6–12 months. This highlights the gap between Anthropic’s current risk assessments and the longer-term concerns raised by Hubinger, notes Seeking Alpha.

This week, OpenAI Chief Scientist Jakub Pachocki warned that the world is not prepared for the rapid advancement of AI and said he expects developers to voluntarily slow down their work in response to these risks, according to Bloomberg.

This article was AI-translated and verified by a human editor

Share

Trending

Stock Screener
Buy
Sell


















Small Caps
Investment and Finance News