Anthropic researcher Jacob Coxon warns of AI risks without global coordination

Editorial illustration: Two hands hold a curved metal barrier around a blue globe, with a glowing golden lattice cube above its northern hemisphere.

In brief

  • Jacob Coxon resigned from Anthropic on September 8 citing AI safety concerns.
  • He appeared on NBC's Meet the Press to argue for global AI lab coordination.
  • Coxon warned catastrophic outcomes could materialize within 6-12 months without oversight.
  • Anthropic's alignment lead estimated >10% probability of AI-driven extinction within a decade.

The resignation and the warning

Coxon accused both OpenAI and Anthropic of treating humanity's safety like an acceptable trade-off in the race to build smarter machines. He stressed that current AI models don't pose an immediate existential threat. The danger, he argued, lies in what could emerge over the next 6 to 12 months as capabilities continue to scale. Coxon warned that catastrophic outcomes from advanced AI could materialize by the end of the decade.

His framing distinguishes between present-day systems and the trajectory ahead. It's not today's models that concern him—it's the velocity of improvement and the absence of coordination mechanisms to manage it.

Anthropic's alignment gap

The critique carries weight partly because of who's making it. Evan Hubinger, Anthropic's alignment science lead, has acknowledged the company lacks a clear strategy for solving alignment problems related to superintelligence. More starkly, Hubinger has publicly estimated a greater than 10% probability that AI could lead to the eradication of humanity within the next decade.

Anthropic has long positioned itself as the "responsible" AI lab, the one that takes safety seriously enough to publish detailed model evaluations and implement usage policies more restrictive than competitors. Yet Coxon's critique suggests the gap between branding and practice may be wider than the company's public-facing materials imply. The company is pushing toward an IPO that could value it at up to $2 trillion—a timeline that may amplify pressure to prioritize growth over caution.

Industry pattern

This isn't isolated. OpenAI has experienced high-profile departures over safety disagreements in recent years, including the dissolution of its superalignment team. The pattern suggests that as AI labs scale, internal safety advocates face mounting friction. Coxon's departure and public statement underscore a structural tension: the labs most capable of building advanced AI systems are also the ones most incentivized to move fast. Both OpenAI and Anthropic have prioritized speed of development over the kind of thorough safety assessments that the technology demands.

The prescription

Coxon's remedy is straightforward: AI labs and government bodies need to collaborate on oversight frameworks before the technology outpaces any ability to govern it. Whether that coordination emerges—or whether competitive pressure overwhelms it—remains an open question.