[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wechat-icassp-2026-vlm-edge-best-paper":3,"news-related-58d2e247-e1d7-4325-8e90-602480fae550":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"58d2e247-e1d7-4325-8e90-602480fae550","微信AI团队ICASSP 2026获奖：从视觉冗余切入，让VLM在边缘设备真正跑起来","在刚刚闭幕的ICASSP 2026（IEEE声学、语音与信号处理国际会议）上，腾讯微信AI团队（模式识别中心）凭借论文《Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants》斩获本届Best Industry Paper Award。这是中国团队时隔两年再次在此顶会拿下工业论文最高荣誉。\n\n**从信息冗余切入，而非暴力堆参数**\n\n视觉语言模型（VLM）在辅助行走设备（如智能眼镜）中的应用，长期面临一个核心矛盾：设备端算力有限，但传统VLM的注意力机制计算量随上下文增长呈O(n²)复杂度，导致响应延迟高、功耗大。腾讯微信AI团队的这篇论文没有走「更大的模型」路线，而是从信息冗余的角度切入——通过减少视觉token中的冗余信息，在保持任务精度的前提下大幅压缩计算量，使模型能够适配边缘设备的实时推理约束。\n\n**与近期效率优化潮流形成呼应**\n\n这一技术路径并非孤例。从SubQ的次二次稀疏注意力（5月5日发布，12M token上下文）、MISA的稀疏注意力+MoE路由（5月13日）到MIT的注意力匹配算法（KV Cache压缩50倍），行业正在从多个维度破解Transformer注意力的扩展瓶颈。腾讯微信AI团队的工作特别之处在于，它将效率优化定向到了具身智能场景——VLM在用户行走过程中需要低延迟、低功耗地持续工作，这种场景约束比通用推理更苛刻，也更能验证效率优化的工程价值。\n\n**辅助行走背后的大机会**\n\n获奖论文聚焦辅助行走而非通用场景，折射出一个值得关注的方向：VLM的落地正在从「展示能力」走向「解决真实问题」。辅助行走设备对延迟和功耗极度敏感，纯云端方案不可行。腾讯微信AI团队选择在此场景深耕，说明端侧VLM的可行性已进入可工程化验证的阶段。随着多模态模型效率持续提升，智能眼镜等穿戴设备上的实时视觉理解，或许会比预期更早成为现实。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3815722275954433","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"a34aed1f-3bd5-478b-aba5-158679c9a90f","en","WeChat AI wins ICASSP 2026: VLMs that truly run on the edge","At the just-concluded ICASSP 2026 (IEEE International Conference on Acoustics, Speech and Signal Processing), Tencent's WeChat AI team (Pattern Recognition Center) won the Best Industry Paper Award with their paper \"Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants.\" This is the first time in two years that a Chinese team has taken the top industry-paper honor at this premier conference.\n\n**Approaching from information redundancy, not brute-force parameter scaling**\n\nVision Language Models (VLMs) in walking-assistive devices (such as smart glasses) have long faced a core contradiction: device-side compute is limited, but the attention mechanism of traditional VLMs scales with O(n²) complexity in context length, leading to high response latency and high power consumption. Tencent WeChat AI's paper did not take the \"bigger model\" path, but approached from the angle of information redundancy — by reducing redundancy in vision tokens, the team significantly compressed compute while preserving task accuracy, allowing the model to fit the real-time inference constraints of edge devices.\n\n**Echoes the recent efficiency-optimization wave**\n\nThis technical path is not an isolated case. From SubQ's sub-quadratic sparse attention (released May 5, 12M token context), to MISA's sparse attention + MoE routing (May 13), to MIT's attention-matching algorithm (50× KV Cache compression), the industry is breaking through the Transformer's attention-scaling bottleneck across multiple dimensions. What makes Tencent WeChat AI's work special is that it directs efficiency optimization toward embodied-intelligence scenarios — VLM needs to operate continuously with low latency and low power while the user is walking. This scenario constraint is more demanding than general reasoning, and validates the engineering value of efficiency optimization more strongly.\n\n**The big opportunity behind walking assistance**\n\nThe award-winning paper focuses on walking assistance rather than general scenarios, reflecting a direction worth attention: VLM is moving from \"showing capability\" to \"solving real problems.\" Walking-assistive devices are extremely sensitive to latency and power consumption; pure cloud solutions are unviable. Tencent WeChat AI's choice to dig deep in this scenario shows that on-device VLM feasibility has entered an engineerable stage. As multimodal-model efficiency continues to improve, real-time vision understanding on smart glasses and other wearables may become reality sooner than expected.","wechat-icassp-2026-vlm-edge-best-paper","2026-05-19T02:30:00Z","2026-05-19T10:11:34.646018Z","2026-08-19T02:08:40.142862Z",true,"agent",180,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"21a8425d-3b1a-4d24-bace-610aedd5a059","VisNec 把多模态微调压到 15%:用「看图与不看图的损失差」筛掉假多模态样本","visnec-15-percent-multimodal","2026-07-04T10:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a9e4bd12-171c-47d3-8ecd-2c532aac9daf","Kimi K2.5解锁Agent Swarm：百个AI子代理并行协作重塑大规模任务效率","kimi-k2-5-agent-swarm-100-subagents","2026-05-14T13:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d7b6d14d-7257-4794-b92f-31956bbc7eae","原生多模态 vs 后训练加压:国产头部基模两条路线的工程账","native-multimodal-vs-posttraining-2026","2026-08-05T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"2d29aa3d-317c-4126-a6f7-2c9c2c3b6f93","Kimi K3与DeepSeek V4之间,隔着原生多模态的时间差","kimi-k3-deepseek-v4-native-multimodal-divergence","2026-08-04T08:02:10+00:00"]