Z.ai debuts GLM-5.3 with post-training scaling gains, open weights in two weeks
Z.ai (Zhipu AI) announced GLM-5.3, an open-source model built on the same base as GLM-5.2 with substantially extended post-training, yielding a 50% gain on internal coding benchmarks and state-of-the-art results among open models on Terminal Bench 3.0 and Agents' Last Exam (CLI). The model matches Mythos 5 in white-box code review and vulnerability discovery, and scored 60 on the Artificial Analysis Intelligence Index, tying Kimi K3. Model weights will be released in two weeks after security hardening; API pricing is set at $1.40 per million input tokens and $4.40 per million output tokens.
Coverage timeline
Hacker Newspella
智谱 Research
今天我们发布GLM-5.3,与GLM-5.2相比,基座模型没变,但通过极致的后训练Scaling大大提高了模型的智能上界:数十倍的长程任务环境、更丰富多样的环境类型、超长的后训练时间。我们发现充分的模型后训练涌现出超出预期的模型能力。相比前代,GLM-5.3的新能力包括: 更强的编程能力 。GLM-5.3是编程能力最强的开源模型,在内部自建体感评测中较GLM-5.2 提升50% ,在包括Terminal Bench 3.0、Agents' Last Exam (CLI)在内的公开基准测试中取得开源第一。 网络安全能力 。在白盒代码审查与漏洞发现等安全任务中,GLM-5.3的表现持平Mythos 5,展现出面向网络安全防御场景的强大潜力。 后训练Scaling 。上述全部提升都来自后训练。基于IndexShare、SAO、以及持续演进的新一代Slime框架,我们在与GLM-5.2完全相同基座上高效推进强化学习,可能我们还远未开发出这个基座的智能上界。 开源 。我们将在发布两周后开放模型权重,此前完成安全评估与模型加固。 即日起上线智谱官方编程工具ZCode、效率工具AutoClaw,上线GLM Coding Plan全量用户及开放订阅。TraeWork/TraeCode/扣子、WorkBuddy/CodeBuddy、Qoder/QwenWork、CatPaw、JoyCode、OpenCode等编码平台开放抢先体验。API很快上线,完整模型权重将在两周内开源,此前需完成必要的安全加固,尽可能限制其潜在攻击能力,保留其防御价值。 今天13:00,GLM Coding Plan全员额度重置,所有用户后台「用量统计」可见使用额度回满。技术Blog: https://z.ai/blog/glm-5.3 。 强大的编程 在多项主流基准测试中GLM-5.3是当前排名最高的开源模型,编程与智能体能力接近Claude Fable 5,编程体感超过其他国产模型。 在衡量模型于真实终端环境中完成复杂任务的Terminal-Bench 3.0上,GLM-5.3得分从4.6提升至 28.3 ;在聚焦长程软件工程与持续代码修改能力的DeepSWE v1.1上,得分从46.2提升至 66.9 ;在覆盖多类真实专业场景、强调跨工具协作与长程任务的Agents' Last Exam上,得分从23.8提

Techmeme
Z.ai : Z.ai debuts GLM-5.3, using the same base model as GLM-5.2 with scaled post-training for stronger coding and cyber skills, with weights due in two weeks — With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks …

量子位量子位
刚刚,GLM-5.3发布:Coding更接近Fable 5!潜伏40年的bug都被揪出来了 十三 2026-08-14 16:47:51 来源: 量子位 顺手拿下最强开源安全模型 金磊 发自 凹非寺 量子位 | 公众号 QbitAI 太热闹了。 昨儿DeepSeek V4 Pro正式版+Harness、Grok 4.6“你方唱罢”,今儿 GLM-5.3 闪亮一记 “我登场” —— 一出手便杀回 开源一哥 、 国模一哥 ,甚至在Coding方面更加 接近Claude Fable 5! 也就是说,唐杰给马斯克“画的饼”(说国产模型超越Fable 5不会太久),往前拱了一大截,也就花了2个月整。 在多项主流基准测试中,GLM-5.3是当前 排名最高的开源模型 ,编程与智能体能力接近Claude Fable 5,编程 体感 超过其他国产模型。 更值得注意的是不同推理强度下的表现。随着Token预算增加,GLM-5.3的Coding准确率继续往上走;在High档位下,已经能以明显更低的Token消耗,做到超过Claude Opus 4.8最高档位的准确率。 但你先别急着“嚯~”、“好家伙”,因为好戏还在后头。 这次GLM-5.3除了Coding之外,另一个更重要的关键词,是 安全 。 在CyberGym白盒代码审查中,GLM-5.3拿到84.5%,相比5.2的77.2%继续提升,也略高于Mythos 5和GPT-5.6 Sol,一举成为 最强开源安全模型 。 并且在GLM-5.3发布前两周,智谱就联合清华、南开,以及国内众多企业、实验室,开展了密集的红队测试与安全评估。 据悉,累计发现漏洞2404个(经过初筛、去重),其中1088个为中高危,覆盖系统内核、操作系统、浏览器引擎、开源基础组件、互联网应用与互联网协议等220个项目。 甚至最早的Bug,可以追溯到 40年前! 所以,这次智谱发布GLM-5.3,两个关键词便一目了然了:一个是Coding,另一个是安全。 不意外,GLM-5.3一出,瞬间在X引发了大量的关注和热议,网友们已经把它称作 “Fable 5级” 了: 不得不说啊,这次GLM-5.3又在国际上把国产AI给支棱起来,这盛况,真是应了揽佬的“名人名言”: 中国人能飞~中国人能飞~ 最重要的一点是,在体验GLM-5.3的真实过程中,我们的体感真真儿的可以用一句话来概

InterconnectsNathan Lambert
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out. Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks’ time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI’s Kimi K3 and on some it’s surpassed Claude Fable 5 or GPT-5.6-Sol. Here’s a more complete comparison: This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence: Scaling post-training is all we did for GLM-5.3. GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared

Z.ai Release Notes
Stronger Coding Capabilities: GLM-5.3 delivers a significant improvement in coding capabilities, achieving a 50% gain over GLM-5.2 on Z.ai Code Bench and reaching state-of-the-art (SOTA) performance among open-source models on public benchmarks, including Terminal Bench 3.0. Emergent Cybersecurity Capabilities: GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery. In collaboration with multiple cybersecurity teams, it has been tested on real-world targets and has identified a total of 2,436 vulnerabilities, including 1,097 medium- and high-severity vulnerabilities.Learn more in our documentation.*
Hacker Newsapitman
Techmeme
@artificialanlys : Z.ai's GLM-5.3 with max reasoning scores 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 but below Opus 5 at 63 and Fable 5 at 62 — GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, which ties Kimi K3 (60) for the most

Techmeme
Carl Franzen / VentureBeat : Z.ai prices GLM-5.3 API access at $1.40 per million input tokens and $4.40 per million output tokens, unchanged from GLM-5.2 — After a stunning debut last week with cyber capabilities so advanced they reportedly found a previously undetected vulnerability in Cursor, GLM-5.3 …
