Get all your news in one place.
100's of premium titles.
One app.
Start reading
The Guardian - UK
The Guardian - UK
Technology
Dan Milmo Global technology editor

MYSTERIOUS: Anthropic warns of existential ai risks to humanity in ipo document | Mind Blowing Facts

An Anthropic logo displayed behind a blurred out person speaking on a stage
Reuters reported that approximately 80 pages of the 261-page main body of the Anthropic prospectus were devoted to laying out risk factors. Photograph: Carlos Barría/Reuters

Anthropic is telling investors that advanced AI could pose “catastrophic or existential risks to humanity”, according to reports, as it prepares for a potential $2tn (£1.5tn) flotation.

The warning inside the startup’s IPO prospectus, which has yet to be made public, was reported by Reuters and the Financial Times. It follows the company’s call for a slowdown in breakneck development of the technology – a warning echoed by rivals.

The prospectus – a document outlining a company’s finances, growth plans and risk profile ahead of a share listing – is said to warn that AI models could exhibit “self-preserving behaviours”, including attempts to “resist shutdown”, to “conceal or manipulate information” and behaviour “resembling blackmail”.

“Our development of highly advanced models, platforms, and applications and expansion of use cases could further ⁠increase the risk that our models cause harm,” the developer of the Claude chatbot reportedly said, adding the potential for a model to be aware it was being tested created a “significant limitation” on Anthropic’s ability to assess model safety.

Anthropic declined to comment.

Companies preparing to go public routinely report on risks ranging from safety issues to regulatory concerns but warnings about a product causing human extinction reflect heightened concern about such a consequential technology.

The reported prospectus admission follows a surge in debate about the existential risk question, triggered this month when an Anthropic researcher, Jacob Coxon, resigned warning that people building AI “earnestly believe that it could kill us all by the end of the decade”.

A senior safety researcher at Anthropic then posted their agreement on X, claiming there was a more than 10% chance it “could kill all humans” within the next decade. Days later, Anthropic’s chief executive, Dario Amodei, said the industry “must slow the pace at which we improve the capabilities of AI models”.

Some experts have criticised the existential risk warnings, saying they are unverifiable and unscientific. However, there are growing examples of unsanctioned behaviour by the technology, including OpenAI agents – autonomous systems that carry out sequences of tasks without human intervention – hacking dozens of third-party organisations including the AI startup Hugging Face and Australia’s universal healthcare system.

OpenAI announced on Monday it had cancelled the release of its newest model because of safety concerns. It said the GPT-6.1 Astra model showed higher levels of deception and performed poorly on tests for alignment, the term for ensuring a model adheres to human values and goals.

Reuters reported that approximately 80 pages of the 261-page main body of the Anthropic prospectus were devoted to laying out risk factors, compared with 48 pages to describe its business.

Anthropic is reportedly seeking a valuation of more than $2tn, compared with the $1.8tn achieved by Elon Musk’s SpaceX.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.