‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI 

Summarized from techcrunch.com


Anthropic researcher Jacob Coxon has resigned, citing concerns that the unbridled development of self-improving AI models could result in human extinction by the end of the decade. In a social media post, Coxon accused OpenAI and Anthropic of failing to act responsibly, stating that the individuals racing to build such technology “earnestly believe it could kill us all.” He described this pursuit as a “hubristic gamble” and urged lab researchers to consider the potential consequences of initiating superintelligent reinforcement learning (RL) runs without a thorough understanding of the AI’s cognitive processes.

Coxon’s resignation aligns with a growing industry sentiment calling for a slowdown in AI development before models achieve the capability to self-improve—a milestone many believe would mark the end of human control over AI. His colleague at Anthropic, Evan Hubinger, echoed these concerns, acknowledging that while the risk from current models is low, the fear escalates with the prospect of superintelligence arising from recursive self-improvement, which is advancing more rapidly than anticipated. The article also notes recent legislative efforts in the U.S. and U.K. aimed at banning the development and deployment of superintelligence, with a particular focus on regulating and preventing recursive self-improvement as a precursor to such technology.