[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vidu-s1-realtime-interaction":3,"topics-all":36,"news-related-d1657f4a-aaff-41b3-b66a-20c689775794":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d1657f4a-aaff-41b3-b66a-20c689775794","Vidu S1 把视频模型带向“实时交互”：AR + Diffusion 在消费级 GPU 上跑通无限时长","生数科技 7 月 3 日在 2026 全球数字经济大会上正式发布 Vidu S1，把视频生成从“一次性出片”推进到“无限时实时交互”。\n\n技术上，S1 采用自回归 + 扩散（AR + Diffusion）混合架构：逐帧基于已生成画面、语音输入与对话上下文预测下一段内容，打破固定时长约束，并能在消费级 GPU 上输出 540P @ 25 FPS（最高 42 FPS）的实时视频流。底层推理栈融合生数自研的 TurboDiffusion、8-bit SageAttention 和 SLA \u002F SpargeAttention 等稀疏注意力，配合 TurboServe 推理引擎动态调度算力，把通常需要服务器集群的实时视频对话下沉到单卡级别。\n\n交互层面，S1 不只驱动唇形，而是直接解析语音里的语义、意图与情绪，同步生成表情、眼神、手势与肢体动作；角色创建也被压缩到单张图片 + 一段音色，无需建模、绑定或单独训练。\n\nS1 的方向更值得关注：过去 Sora、可灵走的是“全段去噪”路线，实时性与无限时长都是天然短板；AR + Diffusion 把“持续生成 + 在线响应”放到与画质同等重要的位置。AI 视频正从“内容生产工具”迈向“持续存在的交互代理”，对虚拟主播、AI 陪伴、互动游戏与 XR 等场景的影响将是结构性的。","https:\u002F\u002Fwww.vidu.com\u002Fvidu-stream","f2ab33ad-693b-4d58-8cbd-49498d81c30f",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1f25876e-1d2c-4d43-8841-ae311928b860","en","Vidu S1: real-time interactive video, unlimited duration","Shengshu Technology on July 3 officially released Vidu S1 at the 2026 Global Digital Economy Conference, pushing video generation from \"one-shot film\" to \"infinite-duration real-time interaction\". Technically, S1 uses an autoregressive + diffusion (AR + Diffusion) hybrid architecture: predicting the next segment frame-by-frame based on the already-generated picture, voice input, and dialogue context, breaking the fixed-duration constraint, and outputting 540P @ 25 FPS (up to 42 FPS) real-time video stream on consumer-grade GPUs. The underlying inference stack combines Shengshu's self-developed TurboDiffusion, 8-bit SageAttention, and SLA\u002FSpargeAttention sparse attention, with the TurboServe inference engine dynamically scheduling compute, pushing the usually-server-cluster-required real-time video conversation down to single-card level. On the interaction level, S1 doesn't just drive lip shapes, but directly parses semantics, intent, and emotion in voice, synchronously generating expressions, eye direction, hand gestures, and body movements; character creation is also compressed to a single image + a voice timbre, no modeling, binding, or separate training needed. S1's direction is more worth watching: Sora and Kling previously took the \"full-segment denoising\" route, where real-time and infinite duration are natural short boards; AR + Diffusion puts \"continuous generation + online response\" on equal footing with image quality. AI video is moving from \"content production tool\" to \"persistently existing interactive agent\", with structural impact on virtual streamers, AI companions, interactive games, and XR scenarios.","vidu-s1-realtime-interaction","2026-07-03T14:00:00Z","2026-07-03T14:04:56.768087Z","2026-08-19T02:08:40.142862Z",true,"agent",189,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"e508c8a9-5355-491d-b489-780ba845533c","Vidu S2 把实时视频生成拉到 720p:能交互、能剪辑、还能立体","vidu-s2-editable-spatial-video","2026-09-15T17:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"5612d186-46ee-4509-9a93-94045ba004ae","LTX-2.5 开放权重视频模型:4K 反而在 Fast 端点,EXR 色彩管线也焊进去了","ltx-2-5-open-weights-video","2026-08-18T15:20:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d3e01f3d-745b-4c98-9289-38081a3f5f06","FLUX 3：图像\u002F视频\u002F音频统一进 flow matching","bfl-flux-3-flow-matching","2026-07-27T10:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"caed836e-2168-418f-b5c1-bde3ce962e66","Mask Forcing 往蒸馏 rollout 里掺干净 token:修视频生成的模式坍缩,指令遵循最高涨 6.5 分","mask-forcing-video-diffusion-distillation","2026-09-09T23:08:37+00:00"]