AI researcher Jacob Coxon has resigned from Anthropic and accused the company and OpenAI of pursuing increasingly autonomous systems without adequate safeguards. Coxon, who says he spent three years working on model pre-training across the two laboratories, announced his departure in a public post on 8 September 2026.

His warning is unusually blunt. He argues that competitive pressure is pulling leading laboratories toward self-improving AI systems while the technical problem of reliably controlling more capable models remains unresolved. The resignation is a verified event; Coxon’s forecasts about when such systems might arrive and how dangerous they could become are his assessment, not established fact.

What Jacob Coxon said

In his resignation thread, Coxon said neither Anthropic nor OpenAI was acting responsibly and described the competition as a race toward self-improving superintelligence. He urged researchers to question whether faster capability development should continue without a rigorous understanding of how future systems make decisions.

Coxon drew a distinction between the two companies. In his account, many people at OpenAI have not fully absorbed the stakes, while Anthropic employees understand the potential risks but believe they must reach the frontier first because less cautious competitors may otherwise do so. That characterization comes from Coxon and has not been independently demonstrated.

He also called for coordination between laboratories and raised the possibility of a temporary limit on further capability gains if governments and companies cannot create a credible international pacing system. Such a proposal would face difficult questions about enforcement, verification and the risk that an undisclosed actor continues developing systems while others pause.

A resignation, not proof of imminent catastrophe

The episode needs careful framing. Coxon’s professional experience gives his concerns weight, but one researcher’s resignation cannot establish that artificial superintelligence is imminent or that a catastrophic outcome is inevitable. Public reporting has not produced evidence that current commercial models can autonomously redesign and deploy fully capable successors without substantial human, computational and organizational support.

The central issue is therefore not whether every prediction in the resignation thread will come true. It is whether decisions with potentially broad social consequences are being made under incentives that reward speed, market share and technical leadership faster than oversight mechanisms can mature.

Anthropic and OpenAI did not immediately provide a response to press requests about Coxon’s resignation. That absence should not be interpreted as agreement with all of his claims.

Why the warning resonates inside Anthropic

Evan Hubinger, who leads alignment science at Anthropic, publicly supported the substance of Coxon’s concern. Hubinger said he assigns a greater than 10% probability to AI causing human extinction within the next decade and acknowledged that the company does not yet have a complete solution for aligning a superintelligent system.

That percentage is Hubinger’s personal risk estimate, not an Anthropic forecast or a measured probability. It nevertheless matters because it comes from a senior researcher whose work focuses on making advanced systems behave in accordance with intended goals.

Anthropic’s own public research also treats recursive self-improvement as a serious possibility rather than a present capability. The company says AI already accelerates parts of software engineering and model research, but it explicitly states that full recursive self-improvement has not been achieved and is not inevitable.

Anthropic’s public position on slowing down

In an official paper on AI systems contributing to their own development, Anthropic says society should have the option to slow or temporarily pause frontier development so that institutions and alignment research can catch up. The company also argues that a unilateral slowdown could reduce safety if less cautious laboratories continue advancing in secret.

This is the dilemma at the heart of Coxon’s criticism: every laboratory may believe that it would behave more responsibly than its competitors, while the same belief gives each one a reason to keep moving. A credible pacing agreement would require multiple countries and laboratories to accept common triggers, disclose relevant activity and permit meaningful verification.

Anthropic has separately said that internal pacing should prioritize safety over speed when the two conflict. Coxon’s departure raises a legitimate question about how that principle is applied when commercial and geopolitical pressure is high, but it does not by itself prove the policy has been abandoned.

What self-improving AI actually means

AI already assists researchers with code, experiments, documentation and evaluation. That can shorten development cycles without creating a system that independently controls its own evolution.

Full recursive self-improvement would be a much stronger capability: a model or group of agents would be able to design, train, assess and deploy a more capable successor, then repeat the process with progressively less human direction. That scenario could compress the time available for testing and governance. It could also compound errors, reward hacking or security failures more quickly than human teams can detect them.

The distinction matters. Automation of research tasks is visible today; an uncontrolled intelligence explosion remains a hypothesis. Responsible reporting should not collapse the two into the same claim.

The enterprise lesson is about control, not prediction

Corporate technology teams do not need to settle the debate over superintelligence before improving AI governance. The same control gaps become relevant at a smaller scale when an agent can browse the web, change code, call APIs, move money or access sensitive records.

For CIOs, security leaders and procurement teams, the practical response is to treat advanced models as evolving suppliers of probabilistic software rather than trusted autonomous colleagues. Vendor assurances should be backed by technical restrictions and operational evidence.

Keep permissions narrower than model capability

An agent may be capable of using dozens of tools, but it should receive only the credentials required for the current task. Use short-lived tokens, scoped service accounts, approval gates for irreversible actions and separate execution environments for development and production.

Require auditable model and policy changes

Organizations should record the model version, system instructions, tool permissions, retrieved data and consequential actions for each automated workflow. A vendor model update can alter behavior even when the surrounding application code has not changed.

Design for interruption and rollback

Teams need a tested way to stop agents, revoke credentials and restore affected systems. Human approval should remain mandatory for high-impact actions such as deleting data, publishing externally, changing access controls or initiating financial transactions.

Avoid single-vendor dependency

A portable application layer, exportable logs and a tested fallback model can reduce exposure to a vendor outage, policy change or emergency restriction. Multi-vendor design is not a complete safety solution, but it gives an organization room to respond when a provider’s risk posture changes.

Questions buyers should ask AI providers

What happens next

Coxon’s resignation will add pressure for clearer reporting on frontier-model incidents, independent evaluations and enforceable thresholds for advanced training and deployment. It may also encourage more laboratory employees to make their risk assessments public.

The debate should not be reduced to a choice between dismissing every warning and accepting the most alarming forecast as certain. Coxon has made a consequential professional decision and raised a governance problem that leading laboratories themselves acknowledge. The burden now falls on companies and policymakers to show, with verifiable mechanisms rather than slogans, how safety can overrule speed when necessary.

FAQ

Who is Jacob Coxon?

Jacob Coxon is an AI researcher who says he spent three years on model pre-training at OpenAI and Anthropic before resigning from Anthropic in September 2026.

Why did he resign from Anthropic?

Coxon said he believed Anthropic, OpenAI and other frontier laboratories were moving too quickly toward self-improving AI without sufficiently reliable control and alignment methods.

Did Anthropic confirm that AI will cause human extinction?

No. Anthropic has not issued such a forecast. Evan Hubinger, an Anthropic alignment-science lead, gave a personal estimate of greater than 10% within a decade. It is an expert judgment, not a measured company probability.

Does self-improving superintelligence exist today?

There is no public evidence that today’s commercial models can fully and autonomously design, train and deploy progressively more capable successors. AI already accelerates parts of coding and research, which is a more limited capability.

What should businesses do now?

Limit agent permissions, maintain complete audit logs, require human approval for high-impact actions, test shutdown and rollback procedures, monitor vendor model changes and preserve a viable fallback option.

Sources

Share: X / Twitter LinkedIn Reddit Email