Alibaba releases Qwen3.8-Max, open weights due next week
Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter MoE model with 95B active parameters, on August 3, 2026. It tops several benchmarks, including Artificial Analysis's Agentic Index (55.4) and OSWorld-Verified (86.1), and its API is now available on QwenCloud with open weights expected next week. The model also powers Alibaba's new agent product 'Qwen Office'.
Coverage timeline
Hacker Newsai2027
QWEN STUDIODISCORD Today, we are officially releasing **Qwen 3.8-Max**, the most capable model in the Qwen family to date. This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week. Built upon the architectural foundation of Qwen 3.5, Qwen 3.8-Max scales to **2.4 trillion** parameters, delivering comprehensive improvements across coding, work, research, and long-horizon tasks. It can not only answer more challenging questions, but also complete complex tasks end-to-end with greater reliability, producing dependable deliverables. * **Qwen3.8-Max** — now available via QwenCloud: * 2.4T parameters (95B active), with open weights releasing next week * comprehensive improvements across coding, work, research, and long-horizon tasks * end-to-end and dependable delivery of complex tasks * Call via API on QwenCloud. ## Coding For a top model, coding today means far more than writing a function on request — it means ta
机器之心新闻资讯
8月3日,阿里巴巴正式发布新一代基座大模型Qwen3.8,总参数量2.4万亿,在编程(Coding)和专业办公(Cowork)方面能力大幅提升。今日放榜的权威三方榜单Arena中,阿里Qwen模型仅次于Anthropic的Claude系列,整体性能处于全球大模型第一梯队。目前,Qwen3.8的API已上线千问AI平台,并接入阿里今日同步推出的Agent产品“千问办公”。Qwen3.8-Max预计下周开源,同时还将开源 Qwen3.8-27B。 今日起,全球开发者都可通过千问AI平台获得Qwen3.8的API服务,国内每百万Tokens输入12元、输出36元,隐式缓存命中仅1.5元,而输入与输出的国际价格仅为Opus5的40%与24%,实现相近智能更高性价比。 图说:权威三方榜单Arena全球排行榜更新 此次发布的是千问大模型系列中尺寸最大、性能最强的旗舰模型Qwen3.8-Max,它支持视觉理解,上下文长度达1M Tokens。在模型架构上,Qwen3.8通过稀疏MoE架构与混合注意力机制的联合优化,首次将千问大模型的总参数量拓展至2.4T,激活95B,获得了更高的推理效率、更快的推理速度,整体性能也迎来新突破:在科研复现评测PaperBench等编程智能体评测中,Qwen3.8-Max较上代模型大幅提升28.2分,以93.0分录得评测新高;在WideSearch、Agent's Last Exam等通用智能体评测中,新模型分别取得81.9分和52.4分,位居前沿。同时,Qwen3.8在指令遵循IF Bench评测中斩获82.8分,科学推理GPQA Diamond取得92.6分,系列通用能力均位居前列;在视觉推理BabyVision评测中,Qwen3.8-Max在无工具条件下取得82.0分,超出部分主流模型近两倍,在评估智能体电脑操作能力的OSWorld-Verified评测中,Qwen3.8-Max以86.1分位居主流模型首位。 编程能力(Coding)是前沿大模型 的 核心能力之一,Qwen3.8实现了自主编程的新突破。 在权威CodeArena榜中,Qwen3.8位居全球第四。2年前的前沿模型能写好函数、补全代码,成为程序员的有力帮手,而如今,Qwen3.8可以从一个空文件夹出发,自主完成一个十数天的真实项目交付,全程无需人干预。只需一句“创建一个自进化的智能

量子位量子位
Artificial Analysis榜单:阿里Qwen3.8Agentic能力得分全球第一 量子位的朋友们 2026-08-06 15:43:56 来源: 量子位 8月6日消息,Artificial Analysis今日正式公布了新一期榜单,阿里Qwen3.8-Max在智能体能力(Agentic Index)排行榜中,超越Claude Opus5、GPT5.6,位列全球第一。 Agentic智能体能力代表AI调用工具解决更复杂问题的潜力,是AI落地各专业场景的关键。这一领域冠军长期被Claude、GPT等垄断,此前中国模型最好的成绩是Kimi K3 的50.1分,阿里Qwen3.8此次以55.4分夺下Artificial Analysis的智能体能力排行榜冠军。 本文由阿里巴巴提供,量子位获授权转载,观点归原作者所有。 版权所有,未经授权不得以任何形式转载及使用,违者必究。

Hacker Newsapitman
# Independent analysis of AI Understand the AI landscape to choose the best model and provider for your use case Update Intelligence Index v4.1.1 Intelligence Index v4.1.1 moves 𝜏³-Banking to v1.0.1 and upgrades the grader for HLE, AA-LCR, and AA-Omniscience to GPT-5.6 Luna (medium)Launch Endpoint Accuracy Index Measuring whether provider endpoints serve the same model quality as the reference Highlights ### Intelligence Artificial Analysis Intelligence Index · Higher is better ### Speed Output tokens per second · Higher is better ### Cost per Task Weighted average cost (USD) per Intelligence Index task · Lower is better Find the right model for your use case Get model recommendations grounded in data from our independent benchmarks and live performance and cost monitoring Intelligence 0 1 2 3 4 5 Speed 0 1 2 3 4 5 Cost 0 1 2 3 4 5 Changelog New article published · 6 Aug Launching v4.1.1 of the Artificial Analysis Intelligence Index Methodology updated · 6 Aug Artificial Analysis Inte