Key points
- Anthropic alignment lead puts extinction odds above 10%
- He says Anthropic has no plan yet for superintelligence
- He posted it as a colleague resigned
- Anthropic's prospectus could be weeks away
An Anthropic executive said publicly this week that the technology his company is building might kill everyone. Evan Hubinger, Anthropic's alignment science lead, put the probability above 10% within the next decade. Unlike the colleague whose resignation prompted the exchange, Hubinger still works at the company.
Hubinger posted on X late Tuesday in response to Jacob Coxon, an Anthropic pretraining researcher who had resigned that evening. Coxon said Anthropic and OpenAI are racing to build systems that nobody will be able to control.
"Jacob is correct here," Hubinger wrote, adding that "we really do earnestly believe AI could kill all humans!" He then put a number on the risk, according to Axios. "I personally think it is >10% within the next decade," he wrote. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
The number is what traveled. But the heavier part is what followed: Hubinger said Anthropic does not yet have a plan to solve alignment for superintelligence and, in his view, is not clearly on track to develop one. That assessment came from the executive who leads the company's alignment science.
It's worth being precise about what Hubinger said. The probability was his personal estimate, posted under his own name on his own account. Anthropic has not published a company estimate of the likelihood of human extinction, and its IPO paperwork is confidential.
Alignment is the industry's term for keeping an AI system doing what people intend. In July, Hubinger was among roughly 1,400 researchers who signed an open letter called "Pacing the Frontier," which urged Washington to create the tools needed to "deliberately pace the frontier of automated AI development," CNBC reported.
Hubinger also distinguished between current and future systems. Today's models pose relatively little risk, he said; his concern is about systems capable of improving themselves, according to Newsweek.
Weeks before the expected prospectus
Anthropic confidentially filed its IPO paperwork with the Securities and Exchange Commission on June 1. "This gives us the option to go public after the SEC completes its review," the company said at the time. The company closed a funding round at a $965 billion valuation in May, and investors have since talked about a listing worth around $2 trillion. Anthropic's prospectus could be weeks away. Reuters reported last week that marketing is expected to begin in mid-October at the earliest and the listing to land days before the November midterms, with the public prospectus not expected until late September. That is a reported timetable rather than a filed one, and the people who described it cautioned that the plans could change.
Most companies say as little as possible in that window. Anthropic's people are saying the loudest possible thing, and the odd part is that it is on brand. Safety is the pitch, and has been since a group of OpenAI staff left to start the company in 2021. Even so, "we do not yet have a plan" is a strange sentence to have sitting in the record this close to a deal. Reuters identified Morgan Stanley, Goldman Sachs, JPMorgan, and Citi as banks preparing it.
None of this is a new position for the company, exactly. In 2023, Dario Amodei and OpenAI's Sam Altman both signed a one-line statement saying "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." What changed this week is the specificity. A named executive, a probability, a decade, and his assessment that Anthropic does not yet know how to solve it.
Anthropic is not alone in saying so. OpenAI chief scientist Jakub Pachocki published an essay called "An Alien Mind" on Sunday. "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," he wrote. "I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established." His employer put that to Washington on Wednesday. OpenAI said it wants mandatory national AI safety requirements, and that until Congress acts it will keep backing state bills, announcing support for four in California, Reuters reported.
Lawmakers noticed. "Safety researchers are resigning, powerful AI models are breaking out of their labs, and companies are racing ahead anyway," Representative Lori Trahan wrote on X on Wednesday. "It's past time for Congress to get off the sidelines and do its job." Trahan introduced the FRONTIER Act in July with Representative Jay Obernolte, a California Republican, and it would set up a framework for governing advanced models. Senator Bernie Sanders and Representative Greg Casar introduced a separate bill this month that would pause advanced AI development until federal safety rules exist. There's still no consensus in Congress on how to regulate the technology, CNBC reported.



