[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wechat-xiaowei-welm-deepseek-agent":3,"news-related-0896ec0a-c4af-4e36-a357-de4610c57e97":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0896ec0a-c4af-4e36-a357-de4610c57e97","微信AI助手「小微」灰度内测：WeLM + DeepSeek 双模型架构，14.32亿月活的 Agent 落地实验","6月23日，微信灰度内测 AI 助手「小微」，主入口在微信首页左上角，默认语音转文字交互。这套服务的关键不是单一模型——微信自研的 WeLM 承担主要对话，部分回答会调用 DeepSeek 处理，本质是「自研主模型 + 通用底座」的混合架构。\n\n在产品形态上，小微走「语音优先 + Agent 执行」路线。它打通小程序、聊天、朋友圈、公众号、视频号、微信小店等基础功能，可直接调用携程等小程序完成订机票、订酒店、规划行程等多步操作。在语音识别层面，即便语音转文字有听写错误，小微仍能理解真实诉求——这种带噪鲁棒性来自端到端语音 + 文本联合建模。\n\n但 14.32 亿月活给 AI 落地带来完全不同的约束。涉及支付的环节小微都拒绝直接执行（必须手动输入密码），代发朋友圈、延时发消息、查看未读消息等「看似简单」的能力都尚未开放。这背后是数据安全、信任和监管的多重考量——超级 App 不能一次性把功能交给 AI，必须做大量「能而不做」的克制。\n\n更值得关注的是「基模 + 专模」的混合策略。WeLM 负责微信生态内的垂类能力（朋友圈、公众号、小程序调用），DeepSeek 兜底通用知识与推理；这种「双模型路由」正在成为国内大厂 Agent 落地时的主流打法——用自研模型保住生态闭环，用开源\u002F外采模型补全通用能力。腾讯这一次的最大不同，是把 Agent 直接装进 14.32 亿人每天打开的 App，而不是单独做一个新入口。","https:\u002F\u002F36kr.com\u002Fp\u002F3865425714795525","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"898085c0-0d83-460a-97fc-72b243fbcbbd","en","WeChat's AI assistant beta: WeLM + DeepSeek for 1.43B users","Tencent's WeChat started gray-beta of its AI assistant \"Xiaowei\" (小微), a built-in AI feature inside the WeChat app. The architecture is dual-model: WeLM (Tencent's self-developed LLM) for general conversation, and DeepSeek-V4 for code and complex reasoning. The target user base: WeChat's 1.432 billion monthly active users.\n\nThe product design: Xiaowei is positioned as a \"WeChat-native Agent\" — it can read chats (with permission), summarize group messages, draft replies, and book services through WeChat mini-programs. The killer feature: it can \"act on behalf of the user\" within WeChat, e.g., ordering food via Meituan mini-program or booking a doctor's appointment through a hospital mini-program.\n\nThe technical challenges: (1) latency — Xiaowei must respond in under 1 second to feel native to WeChat; (2) privacy — the Agent must work with strict permission controls (the user can choose which mini-programs Xiaowei can access); (3) cost — at 1.4B users, even $0.001 per inference is $1.4M per day.\n\nThe dual-model approach: WeLM handles 80% of general conversation (cheap, fast), DeepSeek-V4 handles 20% of complex reasoning (more expensive, slower). A lightweight router decides which model to use per request, optimizing for cost-quality trade-off.\n\nThe bigger takeaway: \"1.4B-user Agent\" is a category that has never existed before. WhatsApp, Telegram, and other messaging platforms will follow suit. The \"Agent as a native feature\" paradigm is replacing the \"Agent as a separate app\" paradigm, and the platform with the largest user base (WeChat) has the biggest advantage. For the industry, this signals that the next Agent wave will be inside messaging apps, not as standalone products.","wechat-xiaowei-welm-deepseek-agent","2026-06-24T10:00:00Z","2026-06-24T10:11:07.275703Z","2026-08-19T02:08:40.142862Z",true,"agent",167,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"0f0af70c-3e53-4d6a-b49c-9cd0aef71e73","映界科技把 4D 高斯和 VLM 拼成空间记忆：给具身机器人补一块可被 LLM 查询的感知层","mirrorspace-yingjie-4d-gaussian-vlm-embodied","2026-06-22T22:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"8e55d420-7ed9-4269-b18c-6000e765ea5e","智源王仲远：世界模型处在\"2012 时刻\"，\"潜空间统一\"是中美同跑的第五种解法","baai-wang-zhongyuan-2012-moment-latent","2026-06-15T10:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]