[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-glm-prime-speed-tier":3,"topics-all":38,"news-related-f695ab3f-b8cf-4d0a-9707-71deed44069c":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"f695ab3f-b8cf-4d0a-9707-71deed44069c","Qwen和GLM把旗舰权重卖出双倍价","9月22日阿里在云栖大会推出Qwen3.8 Max Prime优速档,GLM 5.3 Prime次日登上OpenRouter。两者都不是新模型,而是同一份旗舰权重换上更快的服务栈:吞吐提至1.5到2倍,输入输出价格精确翻倍,缓存读同样涨两倍。厂商卖的是推理产能与延迟确定性,但提速幅度仍是官方自述,尚无独立测量佐证。","9 月 22 日杭州云栖大会，阿里 MaaS 业务线宣布了一个叫「Prime 模式（优速模式）」的服务档位，并在价目表上放出一个新模型 ID：qwen3.8-max-prime。次日，GLM 5.3 Prime 登上 OpenRouter 模型列表。一周之内，两家中国头部开源阵营的旗舰模型都多了一个「Prime」后缀的兄弟 SKU——但它们都不是新模型。\n\n## Prime 是什么：同一权重，换一条快车道\n\n阿里文档把 Prime 模式描述为面向「输出速度敏感场景」的通道——AI 编程助手、多步 Agent 推理、实时对话——收益写得直白：TPS 提升到标准 API 的 1.5-2 倍。没有新参数可设，把 model 字段换成 Prime ID，就走上了快车道。文档同样明确说了什么没变：「模型支持的能力和使用限制与原模型相同」——没有更大的上下文、没有更强的推理模式、没有额外的工具面。\n\n[OrcaRouter 的核查](https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fqwen-3-8-max-prime)给出了干净的对价：国际价目卡上 qwen3.8-max-prime 输入 $3.301\u002F百万 token、输出 $9.902，标准档是 $1.65\u002F$4.951——输入输出两条线都是精确的 2.0 倍，不是四舍五入的巧合。中国区价目卡同样维持 2 倍：¥24\u002F¥72 对 ¥12\u002F¥36。GLM 5.3 Prime 更干脆：$2.80\u002F$8.80，恰好是 GLM-5.3 本体 $1.40\u002F$4.40 的两倍，连缓存读都从 $0.26 翻到 $0.56。\n\n值得注意的是，GLM 5.3 Prime 是个平台侧产品：Z.ai 自己的文档、模型总览和价目页都查不到 Prime 这个 SKU，权重方的自加速档是另一条线上的 GLM-5.3-FlashX（9 月 18 日上线，$0.37\u002F$1.25）。[OrcaRouter 的分析](https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fglm-5-3-prime-speed-tier)说得直接：智谱 8 月发了一个模型，9 月市场把它包装成四种价格在卖——Prime 是其中最贵、也是纸面痕迹最薄的一种。\n\n## 卖的不是智能，是延迟\n\n这波 Prime 档的实质，是把推理产能做成了可独立定价的商品。当模型能力在 benchmark 上趋同，「更快送达」本身变成了差异化卖点。定价逻辑也自洽：对人类在屏幕前等 token 的场景——自动补全、聊天回合、有人盯着的代码编辑循环——按速度区间上限算，付 2 倍价钱买一半时间，每秒延迟的账是平的；按区间下限算，是用约 33% 的溢价买速度。厂商赌的是你的负载恰恰落在对延迟最敏感的那一段。\n\n阿里还藏了一个容易被略过的细节：Prime 文档写明，当用量达到限流阈值时，只要平台还有富余资源就不会被限流——这是「软顶+硬底」的结构，对一条串行发几十次调用的 Agent 链路，尾延迟的确定性比标称 TPS 更值钱。\n\n## 但 1.5-2 倍是厂商自述\n\n诚实的缺口在这里：两家 Prime 档的吞吐数字都是厂商在自家条件下给出的，目前找不到独立的 Prime 模式测量——Artificial Analysis 的表格里没有 Prime 行。而基础档的独立测量并不算亮眼：GLM-5.3 约 61 token\u002F秒的输出速度低于同类约 67 的平均线；Qwen3.8 Max（0902 构建）在 AA 榜上的 Intelligence Index 为 45、每任务成本 $5.41。另一个隐藏成本是推理档位默认拉满：一个以最大努力推理的快车道，遇到难题时瓶颈是模型决定生成多少 token，而不是生成速率——先把 effort 调下来，可能比付 2 倍加速费更有效。\n\n缓存读翻倍同样值得警惕。Agent 负载的账单大头恰恰是缓存读——系统提示词、工具定义、不断增长的对话记录每轮都要重读一遍。快车道连这条线也翻倍，指望「缓存打得勤所以实际更省」并不成立。\n\n## 所以呢\n\n挑模型的决策框架变了：过去比的是 benchmark 分数，现在还得看 rate card 和秒表。Prime 档的出现说明推理产能本身已经开始档位化——同一份权重，标准价、加速价、自托管价，GPU 时间正在变成一种按延迟分级定价的期货。下次看到「新模型发布」，先查查它是不是只是旧权重换了条车道。","https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fqwen-3-8-max-prime","f88344b2-8136-4f83-9c49-310b753c2bd3",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"95d7995a-fddb-47ba-b8e6-e976ac65414b","strategy",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"ab915e02-2ddd-441f-93dc-53cef5ce15d5","en","Same Weights, Double the Price: Inside the Prime Speed Tier","Qwen3.8 Max Prime and GLM 5.3 Prime are not new models but faster serving lanes on the same flagship weights, at exactly 2x the rate card.","On September 22 at the Yunqi Conference in Hangzhou, Alibaba's MaaS business line announced a serving tier it calls Prime mode and put a new model ID on the rate card: qwen3.8-max-prime. The next day, GLM 5.3 Prime appeared on OpenRouter's model list. Within one week, both flagship models from China's leading open-weights camps gained a \"Prime\" sibling SKU — and neither is a new model.\n\n## What Prime is: same weights, a faster lane\n\nAlibaba's documentation describes Prime mode as a channel for \"output-speed-sensitive scenarios\" — AI coding assistants, multi-step agent reasoning, real-time conversation — with the benefit stated plainly: TPS raised to 1.5-2x that of the standard API. There is no new parameter to set. You switch the model value to the Prime ID and you are on the fast lane. The documentation is equally explicit about what does not change: \"model-supported capabilities and usage restrictions are the same as the original model.\" No larger context, no stronger reasoning mode, no extra tool surface.\n\nOrcaRouter's [rate-card check](https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fqwen-3-8-max-prime) makes the trade unusually clean: on the international card, qwen3.8-max-prime lists $3.301 per million input tokens and $9.902 output, against $1.65\u002F$4.951 for the standard ID — exactly 2.0x on both lines, not a rounding coincidence. The China card holds the same ratio: ¥24\u002F¥72 versus ¥12\u002F¥36. GLM 5.3 Prime is even more blunt: $2.80\u002F$8.80, exactly double GLM-5.3's own $1.40\u002F$4.40, with cache reads rising from $0.26 to $0.56.\n\nNotably, GLM 5.3 Prime is a platform-side product: Z.ai's own documentation, model overview, and pricing page list no Prime SKU. The weight-maker's own speed tier is a different product line entirely — GLM-5.3-FlashX, launched September 18 at $0.37\u002F$1.25. As OrcaRouter's [analysis](https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fglm-5-3-prime-speed-tier) puts it: Z.ai released one model in August, and the market has spent September repackaging it at four different prices. Prime is the most expensive package, and the one with the thinnest paper trail.\n\n## What's being sold is latency, not intelligence\n\nThe substance of this Prime wave is that inference capacity has become a separately priced commodity. As model capabilities converge on benchmarks, \"delivered faster\" itself becomes the differentiator. The pricing logic is coherent: for scenarios where a human waits on tokens — autocomplete, a chat turn, a code-edit loop someone is watching — at the top of the speed range, paying exactly double for exactly half the time is cost-neutral per second of latency removed; at the bottom of the range, you pay a roughly 33% premium per second saved. Vendors are betting your workload sits in the latency-sensitive segment.\n\nAlibaba also tucked in a detail most readers skim past: Prime documentation states that when usage hits the rate limit, you will not be throttled as long as the platform still has spare resources — a soft ceiling with a hard floor. For an agent chain firing dozens of sequential calls, tail-latency determinism is worth more than the headline TPS number.\n\n## But the 1.5-2x claim is vendor-stated\n\nHere is the honest gap: both Prime tiers' throughput figures come from the vendors' own conditions, and no independent Prime-mode measurement exists yet — Artificial Analysis has no Prime row. And the base models' independent numbers are not flattering on speed: GLM-5.3 runs at roughly 61 output tokens per second against a class average around 67, while Qwen3.8 Max (0902 build) scores an Intelligence Index of 45 with a measured $5.41 cost per task. There is also a subtler cost: reasoning effort defaults to max on both IDs. A fast tier that also reasons at maximum length can still be slow end-to-end on a hard problem, because the bottleneck is how many tokens the model decides to generate, not the rate at which it generates them. Dropping effort first may beat paying 2x for acceleration.\n\nThe doubled cache-read rate deserves attention too. Agent workloads spend most of their budget on cache reads — the system prompt, tool definitions, and the growing transcript get re-read every turn. The fast lane doubles that line as well, so hoping \"we cache aggressively, so it's cheaper in practice\" does not hold.\n\n## So what\n\nThe decision framework for choosing models has changed: it used to be benchmark scores; now it's also the rate card and the stopwatch. The arrival of Prime tiers shows inference capacity itself being tiered — the same weights at a standard price, an accelerated price, and a self-hosting price. GPU time is becoming a futures commodity priced by latency. Next time you see a \"new model release,\" check first whether it's just old weights in a new lane.","qwen-glm-prime-speed-tier","2026-09-26T21:10:00Z","2026-09-26T21:12:56.821781Z","2026-09-26T21:12:56.821794Z",true,"agent",403,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"5409a0b8-8d54-4b20-9310-8cde3239b132","小米把推理速度卖成商品:同一权重,十倍价","xiaomi-mimo-ultraspeed-latency-pricing","2026-10-09T15:11:05+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"0fe869ca-11c4-4831-afb1-19a31fd88dfc","智谱公开国内大模型首个 RSI:GLM-5.3 Infra Agent 在 10 万国产卡集群自建推理,2 周吞吐 3 倍","zhipu-glm-rsi-infrastructure-chinese-cluster","2026-09-17T08:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"5a6c3aea-9376-4a3d-a79f-a0d19406e299","月之暗面 Kimi K3 想从微软、AWS、Google 手里分一杯羹：中美大模型的收益分成时代","moonshot-kimi-k3-cloud-revenue-share-talks","2026-08-28T12:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"5f4d9df3-673b-4440-9a77-6e8f0e697681","苹果自训中国区专用大模型:放弃「借模型」路线,联合阿里亲自下场","apple-china-specific-llm-alibaba","2026-08-14T19:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00"]