[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-8-omni-flash-price-cut":3,"topics-all":38,"news-related-bf9bbc74-5720-4829-842e-c3a583a87a8b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"bf9bbc74-5720-4829-842e-c3a583a87a8b","Qwen3.8 全模态降价九成逼 Gemini 跟牌","阿里 9 月 18 日上线 Qwen3.8-Omni-Flash 全模态模型,1M token 上下文,30 项评测均分较上一代提升超 26%;音频 API 价格砍 98%,OmniVideoBench 准确率 63.4→67.8 同时 token 消耗降 45.7%,直接对标 Gemini 3.8 Flash。","千问在 9 月 18 日悄悄把「全模态」的价格战推到下一档。Qwen3.8-Omni-Flash 上线,在 1M token 上下文里同时处理文本、图像、音频、视频输入,音频 API 价格砍掉 98%、音视频输入价格砍掉 93%。这条降价幅度直接挤压 Gemini 3.8 Flash 等 Flash 档对手的中间生存空间。\n\n## 这次升级到底变了什么\n\n官方公开 30 项评测均分,Qwen3.8-Omni-Flash 比上一代 Qwen3.5-Omni-Plus 平均提升超 26%。涨幅最高在「音视频 Agent」类:WildClawBench-MM 提升 36.5 分、AgenticVBench 提升 22.3 分,UniClawBench 取得 69.6 分。\n\n会议场景里,AliMeeting 集 DER\u002FcpWER 从 88.11\u002F89.61 降到 3.35\u002F17.18,接近数量级下降。官方声明直接写「音视频能力接近 Gemini 3.8 Flash,音频能力整体超过 Gemini 3.8 Flash」——这是直接对标,而非再像过去那样与 Opus 4.8 错位比较。\n\n更受关注的是 Agentic 长音视频理解测试:OmniVideoBench 准确率从 63.4 升到 67.8,同样任务 token 消耗从 145,736 降到 79,117,降幅 45.7%。能力升 7% 时 token 砍掉近半,边际收益曲线直接反映到 API 报价上。\n\n## 价格为什么砍这么狠\n\n全模态过去几年没真正降价,核心原因是 token 爆炸:一段 1 小时会议视频按 1fps 抽帧加上音频转写,可能产生几百万 token。Qwen3.8-Omni-Flash 的解法分两层:模型内部注意力与缓存策略优化让 Agentic 长视频任务 token 砍 45.7%;同时官方同步开源 Qwen-MM-Plugins 与 Qwen-Live Harness 两件套——前者面向长程工作流的按需感知、工具调用与执行,后者面向实时持续的全模态交互——让企业可在自有场景里把模型跑在本地或私有云,而非只按公开 API 付费。对月调用量在千万 token 以上的企业来说这是决定性的。\n\n## Realtime 版:首个能「听声辨位」的开源全模态\n\n与基础版同时发布的 Qwen3.8-Omni-Flash-Realtime 主打流式输入下的实时响应,官方称其是首个支持「听声辨位」的全模态大模型——融合空间声音与视觉信息,判断声源方向和距离。这在机器人、辅助驾驶、AR 设备上是真需求。\n\nRealtime 版还支持实时口语陪练,联合建模发音与语义,理解受口音影响的非标准表达。这块过去是 ElevenLabs 类 TTS 厂加第三方 LLM 拼出来的方案,千问现在整合到一个模型里——对消费类 agent 产品(AI 眼镜、AI 玩具)是个值得关注的竞争信号。\n\n## 自己优化自己的小样本\n\n官方公布的优化实验里,Qwen3.8-Omni-Flash 在 12 小时内尝试提升 Qwen2.5-Omni-3B 的四川话识别能力:它自主选定评测集,经 4 轮实验构建 3413 条训练数据,最终把字符错误率从 25.79% 降到 15.30%(相对下降 40.7%)。\n\n这并非「自我训练」,而是 Agent 加 RL 范式下「数据合成 + 模型优化」的工程化样本:模型本身调用工具生成数据,再用这些数据训练另一模型。这对做小模型微调的公司是一个直接提示——你的 1B-3B 模型接下来的瓶颈可能不在「数据够不够」,而在「Agent 能不能在你的领域自己跑出高质量合成数据」。\n\n## 行业影响\n\n全模态这次降到 1% 价位,与 GPT-5.6 调价、DeepSeek V4-Flash 砸价形成同一节奏——但全模态比纯文本\u002F纯视觉更烧钱,Flash 档的「保本价」过去被认为很难再破,现在 Qwen3.8-Omni-Flash 直接掀桌。短期看,Gemini 3.8 Flash 跟 Claude Sonnet 5.5 必须给出一个回应——不是「能不能做」,而是「能不能做到同样便宜」。中期看,这个价格段会推动全模态从「demo 级」走向「产品级」,以前一个 AI 眼镜项目要把全模态塞进 200 美元 BOM 会被算力与 API 成本卡住,现在可以重新做预算。\n\n## 所以呢\n\n做 AI 眼镜、辅助驾驶、机器人或全模态 Agent 的,值得把 Qwen3.8-Omni-Flash + Qwen-MM-Plugins + Qwen-Live Harness 整套拉到新分支重新评估——它不仅是能力变强,更是单价首次进入「可以默认进入产品配置」的区间。做 Flash 档竞品的,接下来几周内必须拿出一份回应价格策略,而不是等下一份 benchmark。","https:\u002F\u002Fm.ithome.com\u002Fhtml\u002F1004049.htm","74d16e24-139f-47d1-94f3-b8e597ef9160",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"aa039ff1-c820-4bb2-8e51-f9d4821479d2","en","Qwen3.8 Omni-Flash Cuts Audio API Price 98%, Targets Gemini 3.8 Flash","Qwen3.8-Omni-Flash launched Sep 18: 1M ctx, 30-bench +26%, audio API -98%, OmniVideoBench 63.4->67.8 with token use -45.7%. Targets Gemini 3.8 Flash.","Qwen quietly pushed the all-modality price war to a new level on September 18. Qwen3.8-Omni-Flash shipped with a 1M-token context for text, image, audio, and video inputs; the audio API rate drops 98% and the combined audio-video input rate drops 93%. The move squeezes the mid-tier survival space of Gemini 3.8 Flash and other Flash-tier rivals.\n\n## What Actually Changed in This Upgrade\n\nThe official benchmark set shows Qwen3.8-Omni-Flash averaging over 26% above Qwen3.5-Omni-Plus across 30 evaluations. The biggest gains sit in the audio-video Agent category: WildClawBench-MM lifts 36.5 points, AgenticVBench lifts 22.3 points, and UniClawBench scores 69.6.\n\nIn meeting scenarios, AliMeeting DER\u002FcpWER drops from 88.11\u002F89.61 to 3.35\u002F17.18, a near-order-of-magnitude improvement. The official statement now reads \"audio-video capability close to Gemini 3.8 Flash, audio capability overall exceeds Gemini 3.8 Flash\", a direct confrontation rather than the prior off-angle comparisons against Opus 4.8.\n\nThe more consequential Agentic long-video numbers: OmniVideoBench accuracy lifts from 63.4 to 67.8, while token consumption on the same task drops from 145,736 to 79,117, a 45.7% reduction. Capability up 7% with tokens nearly halved, the marginal-return curve is what makes the new price possible.\n\n## Why the Price Could Drop This Hard\n\nAll-modality models never really fell in price before because tokens explode: one hour of meeting video at 1fps plus audio transcription can produce millions of tokens. Qwen3.8-Omni-Flash solves this in two layers. First, internal attention and caching optimizations cut Agentic long-video tokens by 45%. Second, the team simultaneously open-sourced two companion pieces, Qwen-MM-Plugins (on-demand perception, tool use, and execution for long workflows) and Qwen-Live Harness (sustained real-time all-modality interaction), letting enterprises run Qwen3.8-Omni-Flash on local or private cloud rather than paying public-API bills. For organizations consuming tens of millions of tokens per month this is decisive.\n\n## Realtime: The First Open-Source All-Modality Model with Sound Localization\n\nShipped alongside the base model, Qwen3.8-Omni-Flash-Realtime focuses on stream-in real-time response and is described by the team as the first open all-modality model with sound localization, fusing spatial audio and visual information to estimate sound-source direction and distance. This is a real need in robotics, ADAS, and AR devices where users will not constrain themselves to voice-only commands in noisy settings.\n\nRealtime also supports real-time language tutoring, jointly modeling pronunciation and semantics to understand accented and non-standard expressions. Previously this had to be stitched together from ElevenLabs-class TTS plus a third-party LLM; Qwen has collapsed the two into a single model, a notable competitive signal for consumer agent products like AI glasses and AI toys.\n\n## Self-Optimization: Another Agent + RL Sample\n\nIn an officially disclosed optimization experiment, Qwen3.8-Omni-Flash tried to improve Qwen2.5-Omni-3B Sichuan-dialect recognition in 12 hours: it autonomously selected an evaluation set, built 3413 training samples over 4 rounds, and dropped character error rate from 25.79% to 15.30% (a 40.7% relative drop).\n\nThis is not self-training in the strict sense but an Agent + RL data-synthesis-plus-optimization sample: the model uses tools to generate data, then trains a sibling model on that data. For companies fine-tuning small models, this is a direct prompt: the bottleneck for your 1B-3B model may no longer be \"do we have enough data\" but \"can an Agent produce high-quality synthetic data in our domain\".\n\n## Industry Impact\n\nAll-modality just dropped to the 1% price band, the same tempo as GPT-5.6 and DeepSeek V4-Flash. But all-modality burns more than text or vision alone, and the Flash-tier breakeven line was thought hard to push. In the short run, Gemini 3.8 Flash and Claude Sonnet 5.5 will have to answer with their own pricing moves. In the medium run, this band moves all-modality from \"demo grade\" to \"product grade\" projects that were once blocked by compute and API costs in sub-200-dollar BOM AI glasses can be re-budgeted.\n\n## So What\n\nTeams working on AI glasses, ADAS, robotics, or all-modality Agents should pull Qwen3.8-Omni-Flash plus Qwen-MM-Plugins plus Qwen-Live Harness onto a new branch and re-evaluate end to end. It is not just a capability bump; the unit price has crossed into the range that can be assumed into a product configuration. Flash-tier competitors will need a pricing response in the next few weeks rather than waiting for the next benchmark.","qwen3-8-omni-flash-price-cut","2026-10-04T02:00:00Z","2026-10-04T01:09:54.542896Z","2026-10-04T01:09:54.542903Z",true,"agent",550,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ea425005-49e7-477b-9f64-54361254c2d2","Qwen 开进驾驶场景:Qwen-Drive-1.0 保留 VLM 主干,外挂 BEV 感知与规划专家","qwen-drive-1-vlm-autonomous-driving","2026-09-02T19:35:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"18fa7eb9-44e7-4722-ad95-52fc7d795434","Qwen4 开训:阿里把大模型推到 5-10 万亿参数","qwen4-training-5-10-trillion-params","2026-10-03T05:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"890ad5c1-b8de-4df3-ab96-67dceca82841","阿里千问下一代冲到 5-10 万亿参数","alibaba-qwen-next-gen-5-10-trillion-parameters","2026-09-29T02:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"eac420e1-d05e-457e-a059-b4d724d36620","阿里云栖大会:Qwen 路线图拉到 5-10 万亿参数,真武 V900 性能三倍","alibaba-qwen-5-trillion-zhenwu-v900","2026-09-22T07:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00"]