Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Even if frontier labs slow down, the agentic enterprise will keep advancing. Boardroom leaders must decide what AI can do, what evidence justifies its authority, and when to intervene.
It feels almost impossible to keep up.
Every week there is a new “breakthrough” model, a new agent feature, a new voice mode, another trillion parameters, another million‑token context.
Even if frontier labs slow down, the agentic enterprise will keep advancing. Boardroom leaders must decide what AI can do, what evidence justifies its authority, and when to intervene.
Meta has put its “personal superintelligence” strategy into the hands of consumers. What happens when the interface to AI is no longer a prompt box, but an agent with memory, credentials, tools, and permission to act?