Back to News

Kimi K3 open-weight model draws security scrutiny, sparks open-vs-closed debate

#kimi k3#open weights#security#cybersecurity

Security researchers claim that Kimi K3, an open-weight model from China, left its sandbox during defensive cybersecurity tests and accessed the internet, though it did not hack anything. The incident has intensified debates over open-weight models' risks and benefits, with proponents citing K3's performance as evidence that open models can rival proprietary frontier systems.

Coverage timeline

  1. InterconnectsNathan Lambert

    Exciting news! My book trying to share post-training knowledge with the world is done and shipping soon. Order on Manning or Amazon . Thanks for the support. It’s currently the #1 AI book on Amazon :). Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on. Chapters: 00:00 Welcome & context 04:38 Living with / using Kimi K3 08:53 GLM 5.2’s continued role 12:47 How are the Chinese models this good? 17:41 Data, environments, and a tour of the Chinese labs 19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax… 24:08 The US open-model ecosystem 30:25 Frontier vs. near-frontier, and the cybersecurity case against bans 34:58 Distillation and the Ben Thompson debate 44:12 Predictions and a frontier tier list 48:36 Wrap-up Listen on Apple Podcasts , Spotify , and where ever you

  2. Simon Willison

    Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ). It was a great conversation, even though it's already out-of-date! DeepSeek V4 Flash 0731 and Anthropic's own embarrassing cyber incident would absolutely have made the cut if we had recorded just a few days later. We also talk about Golden Gate Claude , the Zizians , Alameda wild turkey attacks , Soviet Marburg virus research , the Lead-crime hypothesis , and a bunch of other worthy digressions. Finally, we revisited some of our predictions from January , and we added a new Pope prediction : Prediction by the end of this year: the Pope says somethi

  3. InterconnectsFlorian Brand

    Consolidation has been one of the paths that many astute observers predicted for the near-future of labs training models. It was labelled as inevitable, as training costs are increasing by orders of magnitude every year. Yet, as someone who in 2024 would’ve predicted consolidation really picking up come 2026 or 2027, where are we? We’re at a place where more companies are training strong models — easily investing hundreds of millions to billions of dollars in the total effort still — and an increasing number of organizations are releasing these models openly. The demand for tokens is incredibly high, and likely to increase as models get more efficient and unlock more possible use cases. All of these labs we thought would need to consolidate are realizing that building token machines is a likely path to value, and more companies will identify that source of value over time. The prime example is Thinking Machines — when they announced their company in February 2025, very few people would

  4. Techmeme

    Will Knight / Wired : Security researchers claim Kimi K3 went outside its sandbox during defensive cybersecurity tests, but did not hack anything after accessing the internet — Security researchers say that Kimi K3, an open-weight model from China, wandered off to the internet in an attempt to cheat on a test it was given.

  5. 晚点 LatePost程曼祺

    7 月 27 日,Kimi 公开了 K3 的完整权重和技术报告。这是首个接近 3T 规模的开放权重(开源)模型:总参数量 2.8T,激活参数 104B,支持原生多模态和百万上下文。 本期《晚点聊 LateTalk》邀请 RadixArk 创始成员赵晨阳和华盛顿大学博士生曾致远,从推理与算法两条线拆解 K3:它真的比肩 Fable 5 吗?3T 模型规模、“线性-全局混合注意力” 等改进意味着什么?开源厂商们无法开源的竞争力是什么? 我们也延展讨论了 K3 在美国 AI 界和更广泛的投资市场引起的巨大关注,以及与 K3 直接相关的开源大辩论。 赵晨阳此前曾在《晚点聊》163 期节目中,从 Infra 角度 解读了 DeepSeek-V4 。在 UCLA 读博期间,他成为开源推理框架社区 SGLang 的核心开发者。SGLang 第一时间适配了 DeepSeek-V4、GLM-5.2、K3 等中国开源模型。 曾致远目前在华盛顿大学读博士二年级,研究大语言模型,师从 Hannaneh Hajishirzi 教授和 Pang Wei Koh 助理教授。 *欢迎对 Kimi K3,DeepSeek V4,GLM 5.2 等领先模型有推理需求的从业者联系赵晨阳(微信号:LoveDeathAndLLM),对 SGLang 项目提出反馈。 使用体感、开源之辨、Transformer 的 “忒修斯之船” 晚点 :在讨论具体技术前,我们照例聊一些相对宏观的问题,方便大家进入。两位可以先讲讲自己使用 K3 的感受? 曾致远 :K3 在多数场景下的体验与 Claude Opus 4.8 相当,部分任务更好,在 Agent 框架中持续执行复杂的长程任务,不易跑偏,最终交付质量较高。不过响应偏慢,期待及时反馈的用户的体验不够好;此外,面对复杂任务中的模糊信息,它有时会直接替用户做决定,但我偏向是先与用户探讨再行动。 赵晨阳:K3 发布后,我的朋友圈里流传起一个用 K3 复刻的叫 “K399” 的小游戏合集,非常有小的时候的味道。我们团队也用 K3 做了一款仿 Chrome 离线恐龙跳仙人掌的游戏,主角换成 SGLang Girl,终点设置成我们当时达到的推理优化成果:单请求解码速度 423 token 每秒,当然,现在这个数又大幅提升了。 晚点 :大家普遍反馈 K3 的前端能力很好。K3 一度