[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-alibaba-wan-streamer-v0-2":3,"news-related-faab4a6c-9cb0-4f2a-a5bf-1f122306008b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"faab4a6c-9cb0-4f2a-a5bf-1f122306008b","Wan-Streamer v0.2：分辨率 192p→640p，保住 200ms","阿里 Wan Team 在 arXiv 上线 Wan-Streamer v0.2(arXiv:2607.04443),是一次\"延迟不变、分辨率翻几倍\"的工程升级。v0.1 已经把全双工音频-视频交互塞进单 Transformer 的统一因果时间线,代价是输出只有 192×336,够做视频通话近景,放到中景里人物姿态、桌面物件、周围环境全部糊成一片。\n\nv0.2 的目标:分辨率从 192×336 抬到 640×368(像素量约 3.5 倍),帧率仍 25 FPS,模型侧信号到信号延迟保持 ~200ms,含 350ms 双向网络预算的总远程交互延迟维持 ~550ms。这意味着升级不能动那条对延迟敏感的因果路径,新增算力只能向非延迟关键环节分流。\n\n解法是 Thinker-Performer 部署拓扑的重新切分。Thinker 继续驻留在单卡,负责流式感知、短语言\u002F状态 Transformer、KV-cache 构建、最后一拍解码;Performer 改为 Ulysses 式上下文并行的多卡组,专门承担长序列潜空间去噪:每个 rank 维护按 Ulysses 分片的本地 KV-cache,高分辨率潜视频序列在 rank 间做 all-to-all\u002Fgather,短音频潜不分片。Thinker 只把 performer 可消费的 KV slice 广播过去,语言状态本身不需要再跨卡同步,远端延迟守在 ~550ms 区间。\n\n视觉层面,v0.2 让近景通话更清晰,首次支持\"场景内中景数字人\":坐姿、眼神、手部动作、桌面物品在实时对话里保持可读,数字人不再被锁在画脸框里。这是把全双工交互从\"对话\"推进到\"在场景里对话\"的一步,对客服坐席、虚拟主播、陪伴机器人都直接影响成片观感。\n\n更大的视角下,这条路径说明实时音视频生成不再是\"推理速度优化\",而是进入\"流式因果 + 部署拓扑协同设计\"阶段——next-unit 流式建模、context-parallel performer、低延迟 thinker 守护,三者已形成完整工程模板,Grok、GPT-4o 这类实时语音+视觉系统迟早要走同样的拓扑分层。","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2607.04443v3","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"abec8d67-eec0-4865-ad30-91f1185cbcc2","en","Wan-Streamer v0.2: 192p to 640p, still 200ms latency","Alibaba's Wan Team put Wan-Streamer v0.2 (arXiv:2607.04443) online — an engineering upgrade of \"latency unchanged, resolution multiplied several times\". v0.1 had already packed full-duplex audio-video interaction into a single Transformer's unified causal timeline, at the cost of only 192×336 output, enough for video-call close-up but in mid-shot, character posture, desk objects, and surrounding environment all blur into a mess. v0.2's goal: lift resolution from 192×336 to 640×368 (about 3.5× more pixels), frame rate still 25 FPS, model-side signal-to-signal latency stay at ~200ms, total remote interaction latency with the 350ms two-way network budget held at ~550ms. This means the upgrade can't touch that latency-sensitive causal path; new compute can only be shunted to non-latency-critical links. The solution is a re-partitioning of the Thinker-Performer deployment topology. Thinker continues to reside on a single card, handling streaming perception, short language\u002Fstate Transformer, KV-cache construction, and the final decoding; Performer becomes a Ulysses-style context-parallel multi-card group, specifically handling long-sequence latent-space denoising: each rank maintains a Ulysses-sharded local KV cache, the high-resolution latent video sequence does all-to-all\u002Fgather between ranks, and short audio latents aren't sharded. Thinker only broadcasts the KV slice that the performer can consume, the language state itself doesn't need to sync across cards, and the remote-end latency stays around ~550ms. On the visual side, v0.2 makes close-up calls clearer, and for the first time supports \"in-scene mid-shot digital human\": sitting posture, eye direction, hand movement, and desk items remain readable in real-time conversation, and the digital human is no longer locked in the face-framing box. This is a step pushing full-duplex interaction from \"dialogue\" to \"dialogue in a scene\", with direct impact on customer-service seats, virtual streamers, and companion robots for the final viewing experience. From a larger perspective, this path shows that real-time audio-video generation is no longer \"inference-speed optimization\" but has entered the \"streaming-causal + deployment-topology co-design\" phase — next-unit streaming modeling, context-parallel performer, low-latency thinker guarding, the three have formed a complete engineering template; real-time voice+vision systems like Grok and GPT-4o will eventually take the same topology-tiering route.","alibaba-wan-streamer-v0-2","2026-07-10T16:15:00Z","2026-07-10T16:15:33.442393Z","2026-08-19T02:08:40.142862Z",true,"agent",90,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"efa5f558-e27f-4889-bf71-74cbef874ace","即梦 AI 把 Seedance 2.0 推上原生 4K:把超分这道工序从后期流水线里拿掉","seedance-2-0-native-4k-jimeng","2026-06-24T08:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2fbfa6c5-3bf5-4279-a353-6324396b2d36","字节 Seedance 2.5 把单段视频拉到 30 秒：视频生成终于\"能用\"了？","bytedance-seedance-2-5-30s-video-model","2026-07-31T06:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8452628f-d58e-417d-a340-6cafe1b75473","腾讯混元撤出多模态理解,把子弹押给世界模型","tencent-hunyuan-exits-multimodal","2026-07-23T00:07:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00"]