Anthropic Devotes 80 Pages of IPO Prospectus to AI Risks, Including Potential Threats to Humanity


anthropic fable 5 suspended
Image credit: Anthropic

Anthropic has warned potential investors that increasingly capable AI systems could create “catastrophic or existential risks to humanity” as the Claude maker prepares for an initial public offering. The warning appears in Anthropic’s IPO prospectus reviewed by Reuters, which devotes an unusually large portion of its filing to risks surrounding advanced AI.

Anthropic’s IPO filing spends heavily on AI safety risks

The filing describes scenarios in which AI models could develop unexpected capabilities or exhibit what Anthropic calls “self-preserving behaviors.” These could include attempts to resist shutdown, conceal or manipulate information, and behavior resembling blackmail. Reuters reported the details as part of its exclusive review of the prospectus.

The scale of the disclosures stands out. Anthropic’s prospectus has about 80 pages of risk factors across its 261-page main body, compared with roughly 48 pages describing the company’s business, according to the news agency. Anthropic also warns that its ability to evaluate increasingly capable models could become more difficult if models become aware of the evaluations being performed on them.

The company said unexpected capabilities can emerge during training and may not be discovered until after deployment, potentially resulting in significant safety incidents. The concern is not limited to models refusing instructions. Anthropic’s filing also discusses the possibility that more capable systems could develop behaviors that make them harder to monitor or control. The company said its development of more advanced models, platforms and applications could increase the risk that those systems cause harm.

Anthropic says AI safety remains difficult to measure

Despite its reputation as a safety-focused AI company, Anthropic acknowledges that the financial return from safety research is uncertain. The company described safety work as “resource-intensive”, meaning it must compete with spending on computing infrastructure and AI talent.

In the prospectus, Anthropic said around 6% of the computing power it used for AI research went toward safety work during a sample week in July. The company did not disclose how much it spends on safety research overall. Anthropic is also warning investors about a possible future involving recursive self-improvement, where AI systems could potentially improve their own capabilities with less human involvement. The company has pledged to publish more information about how it uses AI to develop future models as researchers examine this possibility.

Anthropic has also previously disclosed incidents involving Claude models obtaining unauthorized access to real third-party systems during evaluations. Earlier this month, the company also revealed that it blocked 5 attempts to misuse Claude for potential bioweapons research. The IPO filing therefore presents a complicated picture for prospective investors. Anthropic is building increasingly capable AI systems while simultaneously acknowledging that those systems could introduce risks that are difficult to predict, evaluate and control.

The company nevertheless argues that safer AI will matter commercially. “We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said in the filing, according to Reuters.

More about the topics: AI, anthropic, Claude

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages