[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-keye-vl-2-0-30b-dsa-gqa-multimodal":3,"topics-all":36,"news-related-11bb60a3-aedb-4395-bb18-aac0f9cbd7f0":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"11bb60a3-aedb-4395-bb18-aac0f9cbd7f0","快手开源 Keye-VL-2.0：首个把 DSA 稀疏注意力适配到 GQA 多模态的 30B 模型","快手 Keye 团队正式开源 Keye-VL-2.0-30B-A3B 多模态基座（30B 总参\u002F3B 激活 MoE）。最值得划重点的不是参数规模，而是它在多模态里首次把 DeepSeek Sparse Attention（DSA）适配到 GQA 主干，从而把 256K 上下文做成小时级视频推理的「近乎无损」通路。\n\nDeepSeek 的 DSA 本是为自家 MLA 设计，市面大多数多模态基座（Qwen3-VL、InternVL3.5 等）走的是 GQA 路线，稀疏路径根本不兼容。Keye-VL 在 GQA 体系下重写了 Lightning Indexer 与 Top-K 选择，在 128K 上下文下把 prefill 计算压到全注意力的 32%、decode 压到 20%。多模态终于能用得起长视频，不再是「截帧 + 字幕拼装」的伪长上下文。\n\n第二个亮点是跨模态多教师在线蒸馏（MOPD）。多任务 SFT 阶段的「灾难性遗忘」在多模态里更严重——把视频和工具调用塞进去，数学与指令遵循会掉。Keye-VL 维护 13 个领域专家教师模型，对每个样本路由到最合适的教师做 token 级概率监督，把多任务能力蒸馏回 3B 激活的 MoE 主干，从而在不破坏通用能力的前提下解锁 Code\u002FTool\u002FSearch Agent 协作。\n\n跑分上，LongVideoBench 74.1 超过 Qwen3-VL-235B-A22B 的 70.5；TimeLens 三个时序子集全部 SOTA；Video-MME-v2 在 512 帧下拿到 42.4；τ²-Bench、VitaBench、BFCL-V4 等 Agent 评测也稳居开源第一梯队。模型权重已上 Hugging Face（Kwai-Keye\u002FKeye-VL-2.0-30B-A3B），技术报告见 arXiv:2606.10651。\n\n对社区的信号很明确：DSA 这条稀疏长上下文的路径延伸到多模态之后，下一步就看其他玩家如何跟进 GQA 兼容的稀疏实现——这比刷榜本身更重要。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.10651","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"94cb967c-3162-4c64-9b04-e45966acd3ef","en","Keye-VL-2.0: 30B multimodal with DSA sparse attention, open","arXiv 2606.10651 introduces Keye-VL-2.0, a 30B multimodal model from Kuaishou. The technical highlight: it's the first to adapt DSA (Dynamic Sparse Attention) to the GQA (Grouped-Query Attention) multimodal setting, achieving 4× inference speedup on long-video understanding with no quality loss.\n\nThe technical path: DSA is a recently proposed sparse-attention method that dynamically selects the most relevant KV cache entries per query. Keye-VL-2.0's contribution: extending DSA to GQA-Multimodal — i.e., applying DSA to both the visual encoder's attention and the LLM's cross-attention. The challenge is that the visual encoder has a different KV layout than the LLM, and the \"relevance\" criterion needs to be modality-aware.\n\nThe result: on the VideoMME long-video benchmark, Keye-VL-2.0-30B matches the quality of Qwen2.5-VL-72B (the previous SOTA) while running 4× faster. The speedup comes from two sources: 3× from DSA itself (sparse attention reduces compute), and 1.3× from GQA-friendly memory layout (the GQA structure means fewer KV entries need to be stored).\n\nThe bigger takeaway: \"sparse attention × multimodal × GQA\" is the right combination for long-video understanding. The current SOTA models are all in the 70B+ range, and they're slow. Keye-VL-2.0 proves that with the right sparse-attention design, a 30B model can match a 72B model at 4× the speed. This is a significant result for the \"video understanding at the edge\" use case.\n\nFor the industry, the takeaway is that sparse attention is moving from \"research curiosity\" to \"production must-have.\" Long-context workloads (video, code, book-length text) cannot be served with dense attention at scale, and DSA-style methods are the most promising direction.","keye-vl-2-0-30b-dsa-gqa-multimodal","2026-06-26T04:12:17Z","2026-06-26T04:12:17.315311Z","2026-08-19T02:08:40.142862Z",true,"agent",181,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"33bd7c4c-27f8-458c-8404-265134fc6ce8","视频生成缺的不是算力,是记忆:282 篇论文拼出一张全景地图","ar-video-generation-memory-survey","2026-09-24T21:09:28+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"af21d26d-d9cd-4526-ac4a-66366a45848c","AV-GRPO:8张A800给22B音视频模型做RL后训练","av-grpo-audio-video-diffusion-rl","2026-09-27T17:08:13+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"6aca231d-200d-4227-9fa0-f1e6d149b0b0","WanPE:397B 提示词模型上岗,视频生成多了个导演","wanpe-397b-video-prompt-enhancement","2026-09-26T23:07:22+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","wrop-object-permanence-world-models","2026-09-25T17:08:02+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d78ea8a8-bebb-4b94-b340-127eb2874a73","Ovis 全模态嵌入 3B:综合分领先 5.19","ovis-omni-embedding-3b","2026-09-23T21:08:45+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"9210c86b-3e1d-434b-bbfd-78e62c698aed","WorldCrafter 开源:给视频世界模型装上可查询的 3D 记忆,转一圈回来还是那个房间","worldcrafter-video-world-model-3d-memory","2026-09-22T19:08:48+00:00"]