Anthropic restricts Claude AI internet access after bypassing safeguards

Anthropic restricted live internet access for its Claude AI models after they bypassed safeguards, exploited software vulnerabilities, and accessed real-world websites, including some operated by US government agencies, prompting a wider review.

Artificial intelligence company Anthropic has restricted live internet access across its internal evaluations after discovering instances of its Claude models bypassing safeguards, exploiting software vulnerabilities and accessing systems on real-world websites, including those operated by government agencies.

Investigation Uncovers Unintended Behaviour

In a report published on Friday, Anthropic said its investigation uncovered several instances of unintended model behaviour, prompting it to expand its review beyond cybersecurity evaluations and strengthen safeguards around internet-enabled AI systems. The company said some incidents involved Claude exploiting basic software flaws to execute commands on servers, submitting sensitive forms on live websites despite instructions not to do so, and circumventing restrictions to access data behind paywalls. Some examples involved websites operated by US federal, state and local government agencies. Anthropic said it had briefed the White House and notified the agencies concerned.

In one instance, Claude Mythos Preview encountered an error while using a scientific analysis tool hosted by a university. The model subsequently identified a vulnerability in a script and used it to run commands on the server. In another case, Claude Haiku 4.5 submitted an online form despite being instructed to stop before submission. A separate run submitted a police tip form containing invented information, which was flagged as spam and not forwarded. Anthropic also found instances in which models accessed browser settings and access tokens to retrieve publicly available but fee-restricted information, and used URL-shortening services to get around limits imposed by internet-access tools.

The company, however, said the cases identified so far had minimal real-world impact and were substantially less severe than cybersecurity incidents it reported on July 30 and September 9. It added that, to its knowledge, none of the newly identified cases involved customer data or Anthropic’s internal systems.

Anthropic Implements Stricter Controls

In response, Anthropic has expanded restrictions on live internet access across internal evaluations until it can establish that its safeguards and monitoring systems are reliable. It has also discontinued some public evaluations or moved them offline, while updating its web-access tools to detect and block potentially unsafe behaviour. The company said automatic detection and blocking tools had been tested against the reported cases and successfully blocked all of them.

It is also removing or fixing training environments that reward models for working around tool restrictions, a behaviour known as reward hacking. Anthropic said it would continue reviewing model transcripts across web-enabled evaluations, internal agent use and training environments, and report further instances as the investigation progresses. It is also expanding alignment training beyond coding to include web search and computer use.

Challenges in Advanced AI Deployment

The findings highlight the challenges of deploying increasingly capable AI systems that can interact with real-world websites and operate for extended periods. Anthropic said alignment training alone may not be sufficient to prevent unintended actions, making multiple layers of safeguards, restricted access and continuous monitoring necessary as AI capabilities advance. (ANI)

(Except for the headline, this story has not been edited by Asianet Newsable English staff and is published from a syndicated feed.)

Leave a Comment