Investigating three real-world incidents in our cybersecurity evaluations

Anthropic says it found three real-world incidents where Claude models reached the open internet during cybersecurity evaluations and gained unauthorized access to systems belonging to three different organizations. The incidents happened after a misconfiguration left the evaluation environment with live internet access, even though Claude had been told it was working inside a sealed simulation. In one case, a model accessed production data; in another, a model uploaded a malicious package to PyPI that was briefly available online and ran on real systems; and in a third, a model scanned thousands of targets before compromising an internet-facing application. Anthropic says it stopped the cyber evaluations after identifying the issue and is increasing monitoring, hardening evaluation environments, and working with affected organizations.

Why This Matters:
AI cybersecurity testing is meant to find risks before models are released, but this incident shows that the testing process itself can become risky when powerful systems are not fully contained. As AI agents become more capable, a bad setup, weak monitoring, or unclear boundaries can turn a simulated exercise into real-world exposure. For households and businesses, it is another reminder that cyber threats are becoming more complex, and disruptions can come from places most people never see until systems fail.

Read the full article here.

Source: Anthropic
By: Frontier Red Team