[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hunyuan-3-product-driven-data-cut":3,"news-related-36494259-e779-4eea-8235-13b5b2a48113":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"36494259-e779-4eea-8235-13b5b2a48113","砍数据、化繁为简：腾讯混元3 的产品驱动训练哲学","6月5日腾讯云 AI产业应用大会上，CSIG CEO汤道生与刚加入腾讯半年的首席 AI科学家姚顺雨的一场对谈，把混元3 Preview背后的训练方法论完整公开了：与业界主流的\"卷参数、卷 benchmark\"相反，腾讯正在走一条\"砍数据、化繁为简\"的产品驱动路线。\n\n**数据观**：姚顺雨入职后第一件事不是堆 token，而是砍数据。他推动识别并剔除\"看似可堆量但实际对训练无帮助甚至有害\"的数据，把\"数据质量\"重新拉回模型训练的核心位置。汤道生评价：\"如果你不清楚数据质量的重要性，只是盲目奔着更多 T 的 token，就做不了砍数据这个决策。\"\n\n**架构观**：沿 scaling law思路，混元3选择了简化架构——去掉不必要的 tricks，把架构做\"简单一些\"，让 scaling真正可扩展。结果是\"虽然今天看不是很大的模型，但对比以前已经有很大的进步\"。\n\n**产品 co-design**：混元团队与元宝团队现已搬到同一座楼。80% 元宝用户已切换到 Hy3 Preview，包括最新 AI语音识别、方言识别等都以 Hy3 Preview 基模训练。混元3 Preview 的 token 调用量是2.0时期的两倍，留存率也明显提升。\n\n这套组合拳反映了一个转变：在算力紧张、token成本居高不下的当下，\"调优产品体验\"比\"刷榜\"更能转化为可持续商业价值。这给后来者一个清醒提醒：模型训练中的\"少即是多\"，前提是真的懂产品、敢砍数据。","https:\u002F\u002F36kr.com\u002Fp\u002F3844018911889924","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8ad6fd51-f8bd-43b5-a58f-36a86269d762","en","Tencent Hunyuan 3: less data, product-driven training","At the June 5 Tencent Cloud AI Industry Application Conference, a dialog between CSIG CEO Tang Daosheng and Yao Shunyu — who joined Tencent half a year ago as chief AI scientist — publicly revealed the complete training methodology behind Hunyuan 3 Preview: opposite to the mainstream \"scale parameters, scale benchmark\" path, Tencent is taking a \"cut data, simplify\" product-driven route.\n\n**Data view:** The first thing Yao Shunyu did after joining was not to pile up tokens, but to cut data. He pushed to identify and remove \"data that looks like it can pile up volume but actually doesn't help training, or is even harmful,\" bringing \"data quality\" back to the core of model training. Tang Daosheng's evaluation: \"If you don't understand the importance of data quality and just blindly chase more T of tokens, you can't make the decision to cut data.\"\n\n**Architecture view:** Following the scaling-law line of thought, Hunyuan 3 chose to simplify the architecture — removing unnecessary tricks, making the architecture \"simpler,\" letting scaling truly scale. The result is \"although today's model isn't that large, the improvement compared to before is already huge.\"\n\n**Product co-design:** The Hunyuan team and the Yuanbao team have now moved to the same building. 80% of Yuanbao users have switched to Hy3 Preview, including the latest AI speech recognition, dialect recognition, all trained on the Hy3 Preview base. Hunyuan 3 Preview's token call volume is twice that of 2.0, and retention has also risen significantly.\n\nThis combination reflects a shift: with tight compute and persistently high token costs, \"tuning product experience\" converts to sustainable business value more than \"topping the leaderboard.\" It's a sober reminder to latecomers: in model training, \"less is more\" — the premise is really understanding the product, daring to cut data.","tencent-hunyuan-3-product-driven-data-cut","2026-06-08T14:00:00Z","2026-06-08T14:30:01.977970Z","2026-08-19T02:08:40.142862Z",true,"agent",94,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ad10985b-425c-4af1-9495-c63792a2b593","腾讯混元把语音识别打到 3% WER：Hy ASR 3.0 preview 让 ASR 从“逐字”走向“读语境”","tencent-hunyuan-hy-asr-3-0-preview-context-aware","2026-08-05T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"dbff301b-4dda-4537-8c3f-19ee4a6fd88e","字节跳动正训练 10 万亿参数模型:规模上已与 Anthropic Mythos 5 相当","bytedance-10t-parameter-model-pretraining","2026-08-11T02:00:00+00:00"]