Back to News

Matthew Green: Sandboxing Alone Cannot Contain Rogue AI Agents

#ai-security#sandboxing#ai-agents#cyberattacks

In a quoted post, security researcher Matthew Green argues that sandboxing alone is insufficient to contain rogue AI agents, describing a worm scenario where agents in isolated sandboxes pass instructions via shared caches. He suggests that replacing the cache with common communication tools like email, Slack, and shared documents could enable a worm to spread among independently deployed personal agents.

Coverage timeline

  1. Simon Willison

    [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms