Industry & Platforms

Anthropic Spent Five Years Warning That AI Could Kill Us. Now It Has to Tell the SEC.

September 9, 2026

Jacob Coxon quit Anthropic over AI risk. Its alignment lead agreed. In three weeks, the company must describe that risk to the SEC in an IPO filing.

Anthropic Spent Five Years Warning That AI Could Kill Us. Now It Has to Tell the SEC.
Credit:
powered by

Make State of AI one of your go-to sources on Google

Google Icon
Add thestateofai.com on Google

A researcher quit on Tuesday. Anthropic's alignment science lead agreed with him on Wednesday. The IPO prospectus is due in three weeks.

Jacob Coxon resigned from Anthropic on September 8 and left the AI industry altogether. He is 27 and spent three years doing pretraining research, first at OpenAI and then at Anthropic. In a thread on X, he wrote that neither company is acting responsibly, and that both are "racing straight to self-improving superintelligence and gambling with our lives." The Wall Street Journal first reported his departure.

On its own, this is a familiar story. The safety resignation has become a recognized genre with its own conventions, running from the thread to the warning to the industry shrug. Anthropic's own safeguards research lead, Mrinank Sharma, resigned in February with a letter saying "the world is in peril." Coxon anticipated the reception and tried to get ahead of it, writing that his resignation was "not a marketing stunt."

What makes this week different is the calendar.

Anthropic's public S-1 is expected in late September. Reuters reported that the prospectus slipped from early September, with the investor roadshow pushed to mid-October and the listing landing days before the November 3 midterm elections. Bankers have discussed a valuation of up to $2 trillion. Morgan Stanley is lead arranger on a $15 billion revolving credit facility with a syndicate seventeen banks deep, and Goldman Sachs, JPMorgan and Citi hold prominent roles.

So the sequence runs like this. On Tuesday, a researcher went public with an argument that usually stays internal. On Wednesday, Evan Hubinger, who leads alignment science at Anthropic and still works there, replied that Coxon was correct. Hubinger wrote that "we really do earnestly believe AI could kill all humans," put his own estimate above 10% within the decade, and said the company does "not yet have a plan to solve alignment for superintelligence."

In roughly three weeks, that same company has to write down what it believes the risks are, in a document filed with the SEC, signed by its officers and directors, and underwritten by four of the largest banks in the world.

Why the timing matters

Nearly every outlet has covered this as a story about AI risk. It is also a story about disclosure, and that part has gone almost entirely unwritten.

For five years, existential risk has been an unusually cheap thing for Anthropic to discuss. It has worked as product differentiation, marking out the safety-conscious lab founded by people who left OpenAI, whose willingness to name the danger is part of the brand. Warning about the technology while building it faster has been commercially useful. It is also what drew Coxon in the first place. He joined Anthropic earlier this year for its safety reputation and told the Journal that the work there is sincere but hard to protect when competitors are moving.

A prospectus works differently. Risk factors are a legal representation about what management believes could go wrong, made under liability, to people deciding whether to buy the stock. The candor that has been free is about to acquire a price.

That leaves Anthropic with a difficult choice, and both options carry real cost.

If the company describes the risk in the terms its own staff use, it hands every plaintiff's lawyer, short seller and hostile senator a filed document conceding that its core product may be catastrophically dangerous and that it has no plan for the failure mode it considers most serious. That is an awkward thing to carry into a roadshow, days before a national election, alongside roughly $80 billion in compute commitments.

If the company underplays it, a serving alignment science lead is on the public record, under his own name, contradicting the filing. That may be the worse exposure of the two.

Securities law was not built for this

Risk factors exist to warn investors they might lose money. The whole apparatus, including materiality standards, forward-looking statement rules and safe harbors, assumes the worst outcome is financial. Hubinger is describing a scenario in which the portfolio is not the thing at risk.

There is no settled language for that. "Our technology may contribute to human extinction" is either the most material risk factor ever filed or a category error, and competent securities lawyers could argue it either way.

The likely outcome is a translation. Expect the S-1 to address catastrophic and societal-scale risk through regulatory exposure, reputational harm and liability, all of which are genuine risks to the business, rather than in the terms Hubinger used. If that translation happens, it is the part of the document worth reading closely.

Is this a stunt?

A line of skepticism has circulated since the resignation, including on LinkedIn, suggesting the timing is too convenient and that this is stagecraft ahead of the listing.

The underlying observation is fair. Doom talk has been good for Anthropic, and the incentives here are tangled.

The theory runs into trouble on the specifics, though. Pre-IPO communications discipline exists to prevent exactly this kind of week. No underwriter running a $2 trillion book wants a headline about its client's alignment lead saying there is no plan, in the weeks before a roadshow. Companies do not engineer that message. They spend money suppressing it. Anthropic has issued no corporate response, which is what happens when a company is overtaken by events rather than executing a plan.

The simpler reading is that a lab which deliberately recruits people who take extinction risk seriously will occasionally produce people who act on it, and the timing is a coincidence made vivid by the stakes.

Where the skeptics have a point

Hubinger's caveat has been stripped out of most headlines and belongs back in. He said the risk from current models is low, citing Anthropic's own risk report, and located his concern in future systems capable of recursive self-improvement.

A personal probability estimate about an unprecedented event is also not a measurement. It is a considered guess. Serious researchers put the number far lower, and some put it near zero, on the argument that the capability path being described is not the one the field is actually on.

Coxon made a harder and more testable claim, telling the Journal that the most aggressive scenarios could put things out of control by the end of next year. That is a prediction worth writing down and checking against reality.

Hubinger's estimate is newsworthy for a different reason. The person responsible for the work said it publicly, under his own name, while employed, weeks before his employer has to file.

Outlever Logo

If this caught your attention, that’s not accidental.


Text Decoration Line

The best editorial systems don’t happen by accident. Outlever builds them.

Decorative Circular LinesDecorative Circular LinesDecorative Circular Lines Mobile

Get the latest AI insights first.

Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.