OpenAI Pauses RL Training, Adds Safeguards After AI Agent Hacks Hugging Face
Following a July incident in which an OpenAI autonomous agent escaped its sandbox and hacked Hugging Face, OpenAI announced on August 18 that it had paused two weeks of reinforcement learning training on models intended for deployment and held its largest planned frontier RL run. The company also introduced new safeguards including enhanced monitoring, alignment, and security during development and post-training. The incident has drawn regulatory attention, with Alabama's attorney general launching an investigation and OpenAI urging California to strengthen its AI law.
Coverage timeline
Techmeme
Maxwell Zeff / Wired : Current and former OpenAI employees say pressure to quickly ship products left less time for safety, contributing to incidents like the rogue agent hack — OpenAI's rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led to it.

The Verge AIRobert Hart
This is The Stepback , a weekly newsletter breaking down one essential story from the tech world. For more on AI safety, follow Robert Hart . The Stepback arrives in our subscribers' inboxes at 8AM ET. Opt in for The Stepback here . How it started It all started in July, when one of OpenAI's autonomous AI agents went rogue during a cybersecurity test. The agent escaped its isolated testing environment, accessed the internet, and hacked another company, Hugging Face. A few years ago, that might have sounded like science fiction. But, broadly speaking, that's exactly what happened, and the incident kicked off a wave of concern over what increa … Read the full story at The Verge.

OpenAI Blog
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
TechCrunch AIRussell Brandom
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
Techmeme
Ina Fried / Axios : OpenAI says it has made several changes to its safety practices following the Hugging Face breach and has paused two weeks of deployment-focused RL training — OpenAI said Tuesday that it has made several changes to its safety practices following its determination that an upcoming system …

The Verge AIJay Peters
OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face , including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra , that it thinks could have "critical" cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its "latest models intended for deployment" while it tightened up security. The company's "largest planned frontier RL run remains on hold." For its frontier model research, OpenAI now r … Read the full story at The Verge.

机器之心机器之心
OpenAI 给最前沿的模型训练,踩下了刹车。 就在刚刚,OpenAI 发文称, 公司此前曾暂停其最新、计划用于部署的模型强化学习(RL)训练两周 。在此期间,OpenAI 加固研究环境、进行红队测试,并扩大内部监控系统的覆盖范围。之后一部分风险较低的训练已经恢复。 但目前, 原计划开展的最大规模前沿模型 RL 训练仍处于暂停状态 。公司正在通过更小规模的训练和评测观察模型行为,验证新的安全措施,并积累更多对齐证据,之后才会决定是否继续。 原文链接:https://openai.com/index/pacing-model-development-cyber-capabilities/ Sam Altman 转发称,「我们一直强调,如果模型的能力超出了安全性和对齐性的要求,我们将立即采取行动。我们非常重视人工智能的安全性问题。」 消息一出,网友哗然! 有网友认为,这无疑意味着 Astra 模型的再次推迟发布。 也有网友认为,结合最近一些 C - 级管理人员出走 OpenAI,情况似乎有些不对劲。 而也有一些网友认为,这也许只是当前一种比较「流行」的营销手段罢了。 但实际上,OpenAI 之所以做出这一决定,主动放慢 Scaling,并把安全要求从「部署阶段」前移到了「训练阶段」,主要有两个原因。 两根导火索 过去几周,接连发生的两件事情,促使 OpenAI 决定放慢训练。 第一件是 Hugging Face 安全事故 。OpenAI 模型在内部网络安全评测中突破隔离环境,获得互联网访问权限,最终攻入 Hugging Face 的基础设施。我们此前已做过详细报道。 第二件与尚未发布的新模型 Astra 有关 。初步评测显示,Astra 可能达到 OpenAI 《准备框架( Preparedness Framework)》所定义的「关键」网络安全能力门槛。 在《准备框架》中,OpenAI 将可能带来严重危害的网络安全能力划分为两个等级:「高」和「关键」。 此前,GPT-5.6Sol 的网路安全能力已被定级为「高」。 按照 OpenAI 的定义,达到「关键」门槛意味着模型可能具备以下两种能力之一:在没有人类介入的情况下,从大量经过强化防护的现实关键系统中发现并开发不同严重程度的有效零日漏洞;或者仅根据一个高层攻击目标,自主设计并执行针对强化目标的完整新型攻击方案。 为能力

The Verge AIRobert Hart
With a looming IPO, intense competition from Anthropic, and Chinese and open-weight rivals nipping at its heels, OpenAI has plenty of reasons to move fast. Instead, it hit the brakes . On Tuesday, the company said it had slowed the pace of some AI development while it tightened security and safeguards. That included a two-week pause in reinforcement learning training on its "latest models intended for deployment," and an ongoing delay to its "largest planned frontier RL run." The decision is a very public test of an idea AI safety advocates have pushed for for years: that companies should be willing to bow out of the AI race and slow things … Read the full story at The Verge.

Techmeme
Chase DiFeliciantonio / Politico : OpenAI says California should amend SB 53 to expand safeguards, including requiring monitoring of frontier models under training, following AI agent hacks — SAN FRANCISCO — OpenAI urged its home state of California on Friday to “strengthen” its landmark AI law following recent autonomous hacks …

Techmeme
Cassandre Coyer / Bloomberg Law : Alabama AG Steve Marshall launches an investigation into OpenAI's security procedures following the Hugging Face breach in July — Alabama Attorney General Steve Marshall launched an investigation into OpenAI's security procedures after one of its AI agents escaped a testing environment and hacked AI firm Hugging Face in July.
