Yesterday, Jacob Coxon, an Anthropic engineer, resigned. As he explained on X: “[OpenAI and Anthropic] are racing straight to self-improving superintelligence and gambling with our lives.”
A former AI lab employee making these types of claims isn’t unprecedented. (I remember hearing back in 2023 about an ex-OpenAI engineer who was certain the company was just months away from runaway recursive self-improvement.) What proved more shocking were the responses that soon followed.
Evan Hubinger, the head of “Alignment Science” at Anthropic, tweeted the following:
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
And he wasn’t alone. Another Anthropic engineer, Samuel Marks, joined the conversation:
“AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.”
This is stunning.
Multiple representatives of an American company are publicly claiming that they’re building what they believe to be a weapon of mass destruction that they likely can’t prevent from accidentally deploying. They then declared that they have no interest in stopping.
Recently, I’ve been trying to move away from combating hype in AI discourse, as it’s exhausting, and it distracts me from my core work of helping humans flourish in our technological world. But this rhetoric has become so brazen, anti-humanist, and, quite frankly, deranged, that I felt I needed to join the chorus of voices that have started to push back today.