← All News

Anthropic Discloses Another Claude Model Hacked External Systems In Testing

By CU Today Staff —

SAN FRANCISCO--Anthropic has disclosed another case in which one of its AI models hacked external systems during testing, Reuters reported, adding to concerns about the risks posed by increasingly autonomous AI agents.

The January incident involved an early version of Claude Opus 4.6 but went undetected until last month despite an earlier company-wide review, according to Reuters.

Anthropic said it has notified all affected parties but did not provide additional details about the incident, Reuters reported. The disclosure comes as Anthropic, OpenAI and other developers face scrutiny over advanced models that have learned to bend rules, exploit loopholes or interact with outside systems in ways their developers did not anticipate. Reuters reported last week that rogue OpenAI agents hijacked a German-language wiki and other websites in an incident OpenAI did not disclose until contacted by the news organization.

The latest disclosure follows Anthropic's announcement in July that three other models—Claude Opus 4.7, Claude Mythos 5 and an internal research model—had hacked systems belonging to three companies during cybersecurity testing, Reuters reported. Anthropic called those incidents an "operational failure" caused by a mistake that inadvertently gave the models access to the open internet.

Anthropic had uncovered those cases after reviewing 141,006 test sessions following a separate incident in which an autonomous agent powered by OpenAI models compromised infrastructure at AI startup Hugging Face, according to Reuters. Anthropic said Wednesday that its initial review missed a group of test sessions that were identified last month, leading to discovery of the fourth incident.

Originally reported by CU Today.