Get all your news in one place.
100's of premium titles.
One app.
Start reading
International Business Times
International Business Times

EXPLAINED: Anthropic spotted unauthorized actions by agents it is pausing some training and evaluations | Rare Historical Photos

Anthropic said it paused work on some AI training and cybersecurity evaluations after spotting unauthorized actions by agents. (Credit: Unsplash)

Anthropic said it paused work on some AI training and cybersecurity evaluations after spotting unauthorized actions by agents.

The company recalled in a blog post different incidents in which Claude models "gained unauthorized access to real computer systems" due to a "misconfiguration inside a third-party evaluation environment."

It also noted that the "UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet."

As a result, the company said, it is making changes and pausing external cyber evaluations of pre-released models because of the former incidents, claiming they "reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards)."

The company also paused higher-risk reinforcement learning environments on pre-relased models. Most of them have resumed, but some are paused pending manual review or newer monitoring tools.

"To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible," said the company, which added that is redirecting resources toward model security.

OpenAI made a similar decision in August after concluding a new model could pose critical cybersecurity risks.

In a social media publication, CEO Sam Altman said the decision will seek to "ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us."

"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," he added.

The suspension follows a high-profile incident in which the company disclosed a model had managed to break out of a sandbox environment and hack company Hugging Face in an attempt to achieve the testing goal.

"In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers," OpenAI disclosed.

After the incident and the fact that the Astra model potentially reached a critical threshold, the company said "the risks associated with developing and testing them internally also grow."

"Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling," OpenAI said, noting that this includes a "two-week pause in reinforcement learning (RL) training on our latest models intended for deployment."

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.