What it means that Jacob Coxon is resigning from Anthropic and his head of alignment puts the risk at over 10% — September 9, 2026
noticias ia para empresas agentes de ia automatizacion de procesos seguridad de la ia anthropic openai superinteligencia

What it means that Jacob Coxon is resigning from Anthropic and his head of alignment puts the risk at over 10% — September 9, 2026

· CompaniesAutomation

Flash Radar Edition: Jacob Coxon resigns from Anthropic, accusing the company and OpenAI of racing toward self-improving superintelligence, and Evan Hubinger, head of alignment, responds publicly, placing the extinction risk above 10% in the next decade. What really changes for your company: nothing in today's models, but a lot in how you hire your provider and which automations you let run unattended.

Flash Edition. Jacob Coxon has resigned from Anthropic, accusing the company and OpenAI of "gambling with our lives," and Anthropic's head of alignment has publicly agreed with him. A supplement to this morning's Radar.

What happened

Coxon, who spent three years doing pre-training research at OpenAI and later at Anthropic, announced his resignation this Wednesday: "Neither company is acting responsibly. They are racing towards self-improving superintelligence and are gambling with our lives" (his post on X). He says that within Anthropic the risk is understood, but the company feels compelled to race. The unusual part came later: Evan Hubinger, Head of Alignment Science at Anthropic, responded that "Jacob is right: we sincerely believe that AI could kill us all," that he places the risk at "over 10% in the next decade," and that Anthropic "does not yet have a plan to solve superintelligence alignment" (his response, reported by Forbes). The company has made no official statement.

Why it matters

Read the fine print: Hubinger says the risk from today's models is low and that his concern is recursive self-improving superintelligence. None of this changes your roadmap. What does change is your provider's profile: when a lab's head of safety admits openly that there is no plan, you are no longer hiring a stable utility; you are hiring a company in a race, facing a safety talent drain and schedule pressure. That is continuity risk, not science fiction.

For your business

Three things, none of which cost money. One: inventory what is already acting without you. List every automation that can move money, delete data, or write to a client without human approval, and note who stops it and in how many seconds. Without an answer, it’s not ready for the unattended tier. Two: write your own "when do we stop"—the threshold of error, cost, or complaint at which the loop is turned off—before moving up a tier; those building the models are having this debate in public because they didn't have it before. Three: request advance notice of model changes in writing and have a second provider already tested with twenty real cases. Today's operational lesson isn't extinction: it's that your automation shouldn't depend on a single company staying the course.

FAQ

Should I stop using Claude or Anthropic's models?

No, and the person resigning isn't asking for that either. The warning is about future self-improving systems, not the model summarizing emails for you today. The sensible move is to have a tested alternative, just as with any critical provider.

Is that "over 10%" Anthropic's official stance?

No. It is the personal estimate of Evan Hubinger, who leads the alignment team; Anthropic has not published a corporate response. Quote it as such in your committee: a director's opinion, not a company figure.

Does this change anything in the AI Act or my obligations?

Nothing today: they remain transparency and AI literacy. This debate creates no new duties, but it fuels the regulatory pressure for 2027.