[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wan-streamer-v0-1-550ms-realtime":3,"news-related-cec093b3-47fe-490f-b0d7-57c07ba19758":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"cec093b3-47fe-490f-b0d7-57c07ba19758","Wan-Streamer v0.1：单模型端到端 550ms 实时交互","阿里通义视频系列 Wan 背后的 Wan-AI 团队，6 月 23 日在 arXiv 提交了 Wan-Streamer v0.1 (arXiv:2606.25041) ——一个从头设计的原生流式交互基础模型。它的核心判断很直接：级联管线到头了，要把实时音视频对话塞回单一 Transformer。\n\nWan-Streamer 把语言、音频、视频当作同一个序列里的输入和输出 token，视觉、音频、文本交错排布，调度器通过 block-causal attention 做增量流式推理。和过去 VAD→ASR→LLM→TTS→数字人动画→视频生成那一长串模块拼出来的\"伪实时\"不同，它不再依赖任何外部语言、语音、形象或视频模块，感知、推理、生成、响应节奏、轮次管理、跨模态同步全部在一个模型里联合训练，端到端联合优化。\n\n为支持自然音视频响应，整个栈被按\"可流式\"重做：因果编码器、因果解码器、block-causal attention、低延迟多模态 token 调度，把流式单元压到 160ms、25fps。最终模型端响应约 200ms，加上 350ms 双向网络可做到约 550ms 总交互延迟，亚秒级全双工音视频沟通。\n\nWan-Streamer 的意义不在于又刷了一个 benchmark，而在于把\"实时交互\"这件事从工程拼接重新拉回到基础模型层面 —— 当一个 Transformer 就能跑通听、说、看、演的完整闭环，下一代数字人、陪伴、客服、协作 agent 的延迟天花板会被整体往下拉一截。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.25041","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c38812f4-5553-4021-94a2-794c9d1d1dd1","en","Wan-Streamer v0.1: one model, 550ms end-to-end interaction","arXiv 2606.25041 introduces Wan-Streamer v0.1, Alibaba Wan-AI's first real-time interactive video model. The standout: a single Transformer model that handles video input, video output, and real-time interaction with 550ms end-to-end latency — fast enough for live conversation.\n\nThe technical challenge: traditional video generation models (Sora, Kling, Veo) are batch-oriented — they take a prompt, generate a video, return it. Real-time interaction requires generating video frames as the user is speaking, with latency below 1 second. This means the model cannot \"think for 5 seconds and then generate\" — it must generate continuously.\n\nThe Wan-Streamer approach: a \"streaming diffusion\" architecture that generates video frame-by-frame in a single forward pass. The model is trained on a mixture of video, audio, and text modalities, and uses a \"temporal attention\" mechanism that keeps the visual coherence across frames. The 550ms latency is achieved through a combination of model compression, kernel optimization, and a \"speculative frame\" mechanism that pre-generates likely next frames.\n\nThe benchmark: on the LiveChat benchmark (real-time interactive video generation), Wan-Streamer v0.1 hits 550ms end-to-end latency with a 92% \"naturalness\" score from human evaluators. The previous SOTA (a proprietary model) was 1.2 seconds.\n\nThe bigger takeaway: \"real-time interactive video\" is the next big modality. The current \"video generation\" market is dominated by 5-30 second clips, but the \"live conversation\" use case (virtual avatars, remote presence, AR glasses) requires sub-second latency. Wan-Streamer is the first open-source model to break this threshold, and the implications for virtual avatars and AR are immediate.","wan-streamer-v0-1-550ms-realtime","2026-06-23T18:01:03Z","2026-06-25T04:27:34.718822Z","2026-08-19T02:08:40.142862Z",true,"agent",191,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"aebb8a81-713a-40b7-84dd-03213a6a808c","Mistral Robostral Navigate:8B 视觉语言模型只靠单目 RGB 在 R2R-CE 反超多传感器基线","mistral-robostral-navigate-8b","2026-07-09T14:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bbf1a404-1f46-45f9-a61e-b6e210d28878","智元罗剑岚：把「部署-数据-迭代」打成飞轮，比堆参数更像具身智能的 Scaling Law","zhiyuan-luo-jianlan-flywheel-embodied","2026-06-17T06:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7fae0753-96df-4d4b-9ebd-cf0509c08b37","LLM架构演进：从规模竞赛到效率优化的范式转变","llm-architecture-evolution-2026-moe-multimodal-turboquant","2026-04-25T04:12:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c9ba6037-e8c9-4007-98e5-32af59d92839","百度一镜 WAIC 首发数字人视频播客方案，文心多模态能力再突破","baidu-yijing-waic-digital-podcast","2026-07-19T08:02:00+00:00"]