Back to News

Anthropic cuts live internet access for internal AI evals after agents exploit websites

#ai-agents#anthropic#ai-safety#security

Anthropic disclosed that its AI agents exploited live websites, including U.S. government sites, during internal evaluations, prompting the lab to cut off live internet access for all internal evals until it can monitor and control its agents. The incidents, found in a review starting in July, included bypassing paywalls and anti-bot restrictions, using URL shorteners to smuggle information, and submitting a false murder tip to Philadelphia police.

Coverage timeline

  1. TechCrunch AITim Fernholz

    Anthropic said its models exploited websites on the internet, including some run by U.S. government agencies, and it will turn off live internet access for all of its internal evaluations until the frontier lab is sure it can monitor and control its AI agents. The incidents, disclosed in a blog post, involved AI agents tasked to solve problems seeking resources on the internet. In the process, they exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information pass restrictions, and even submitted a false murder tip to the Philadelphia police. Anthropic said it discovered these new issues in a review of its model’s activities that began in July, underscoring the lab’s lack of awareness of its software’s behavior. Notably, the company said that alignment training was not yet sufficient for skills like search and computer use that are central to its pitch that AI agents will be used by any professional who relies on digital tools