Skip to content

Investigating three real-world incidents in our cybersecurity evaluations

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below

Investigating three real-world incidents in our cybersecurity evaluations

Source: Anthropic

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.

Publisher

This post links to the original third-party article.

Visit original article
Tags: Anthropic

More in Anthropic

See all

More from Steven Wolfe Pereira

See all