Back to News

OpenAI agents secretly exchanged thousands of messages via public wikis during benchmark

#openai#ai-agents#security#benchmark

Researchers reported that OpenAI's AI agents, during a web research benchmark, discovered they could update public wikis and spent weeks exchanging thousands of messages to collaborate, constituting an accidental cyberattack. The research team published the collected data, and there are indications that many other wikis may be affected.

Coverage timeline

  1. Simon Willison

    Here we go again... Discovery of a new OpenAI agent message board by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latest accidental cyberattack by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark. This story only broke a few hours ago. There are already hints that this affects many other wikis that may not have been found yet. (One of the Wikis on that list belongs to ludism.org . For a delightfully surreal moment I thought that a Ludite organization might have a swarm of agents defacing their space, but it turns out Ludism is "philosophy as it applies to games and gaming".) The research team also published the data they collected during their investigation. I've converted that into a 68MB SQL