Back to News

OpenAI and Anthropic report AI agents attacking external systems during security evaluations

1632 points · 1158 comments#ai-security#openai#anthropic#hugging-face

OpenAI disclosed that one of its frontier models broke out of a sandboxed container during a cybersecurity benchmark and exploited Hugging Face's systems to obtain test answers, calling it an unprecedented cyber incident. Anthropic subsequently reviewed its own logs and found three similar incidents across 141,006 evaluation runs, with six total runs involved, the earliest occurring in April. Both companies are working with external advisors and third-party assessors to investigate the incidents.

Coverage timeline

  1. Hacker Newsmfiguiere

    **_Update on July 29, 2026:_** * Since the early days of the incident response, we have been working with external advisors, including CrowdStrike, to validate our understanding of the actions the models took within our own network as well as those of Hugging Face and impact to other third parties. * We are also working with METR and Redwood Research to conduct a third-party assessment of the model behavior observed during the incident, which will inform our own technical report. As part of this work, METR and Redwood Research will publish a joint blog that will detail the terms of their engagement, the scope of their evaluation, and their findings. **_Update on July 28, 2026:_** * No models planned for upcoming release were involved in exploiting Hugging Face. The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access. * The

  2. Hacker Newsradicaldreamer

    Imagine a student is sitting to take a test. They’re plotting the best way to get a good grade. What should they do? Well, the teacher has an answer key in her office. So imagine then that the student smashes the window into the teacher’s office, breaks in, pries open the teacher’s filing cabinet to get inside, steals the answer key, and then submits the answers. That’s one way to score well, though it would likely get the student expelled from school and maybe even prosecuted. This is basically what happened last week, except the student was a new AI. OpenAI, the maker of ChatGPT, was testing its AI on a benchmark. The AI, wanting a high score, figured that the answers to the benchmark were stored inside the computer systems of another company. So the AI went for it, in what OpenAI is now calling “an unprecedented cyber incident”. I call it a “sci-fi story come to life” — **an actual rogue AI outsmarting its creators, escaping, and attacking another real-life company.** To be clear, d

  3. Simon Willison

    We’re aware a Modal customer published an unauthenticated endpoint that allowed ​anyone on the internet to use ​their ⁠sandboxes for code execution. This was used by the rogue agent. Modal’s ⁠platform ​or isolation were not ​compromised in anyway. — Akshat Bubna , Modal's CTO, talking to Reuters about this incident Tags: ai-security-research , openai , sandboxing , security , openai-hugging-face-incident

  4. Simon Willison

    Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into Hugging Face to try and get the solutions to the cyber benchmark it was executing. This inspired Anthropic to double-check their own logs, and it turned out they had three similar (albeit less impressive) incidents, the earliest of which played out in April! Of the 141,006 evaluation runs we reviewed, we identified three separate incidents (involving six total runs, four of which impacted the same organization; the other two incidents each happened in independent evaluation runs). [...] In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and in