[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-xmax-x2-realtime-video":3,"news-related-c463258a-094c-4217-8130-eef2bfacf78c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c463258a-094c-4217-8130-eef2bfacf78c","Xmax X2.0 把实时交互视频模型端侧化:逐帧自回归 + 消费级显卡跑 960p@24fps","7 月 15 日,清华系创立的 Xmax AI 发布通用实时交互 AI 视频模型 Xmax X2.0,并同步开放 API。和 Vidu S1、阿里 Wan-Streamer 等云端实时路线不同,X2.0 把实时与端侧同时塞进了同一份模型权重。核心技术是逐帧自回归生成架构——传统扩散模型需等整段算完再吐结果,X2.0 改为每生成一帧即推送给前端,把批渲染拆成流式输出。配合推理优化与模型压缩,响应压到毫秒级近乎零延迟,分辨率从 X1 的 480p 提到 960p@24fps,延迟与画质同步反向升级。功能层面,X2.0 把 CharX 实时换人、ClothX 实时换装、VibeX 风格迁移、MoX 触屏交互整合到统一架构,镜头流、手势、prompt 都能作控制信号。更值得关注的是端侧落地。X2.0 已在最新款 iPhone 上跑通 384@16fps 本地流式推理,在消费级显卡上稳定运行。实时交互视频从需要数据中心 GPU 集群支撑的能力,推进到手机、智能眼镜随手可用的工具层。这给文旅数字导览、电商实时试穿、短剧互动、AR 滤镜补上了低延迟与终端可部署的关键拼图。但 X2.0 更偏应用加工程成果而非前沿架构突破:逐帧自回归不是新范式,真正考验其长期价值是端侧画质与一致性,以及 SDK 在生产环境的稳定性。这条路要跑通,还得看后续业务场景的真实反馈。","https:\u002F\u002Fplatform.xmax.ai\u002F","3ea560bc-8202-436b-ab01-342a19cc0b4b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"375e22a6-8a75-4846-93ea-31923b1adf36","en","Xmax X2.0: real-time interactive video on consumer GPUs","On July 15, Tsinghua-spinout Xmax AI released its general-purpose real-time interactive AI video model Xmax X2.0 and simultaneously opened its API. Unlike cloud-only real-time routes such as Vidu S1 and Alibaba's Wan-Streamer, X2.0 puts real-time and on-device into the same model weights. The core technology is a per-frame autoregressive generation architecture — traditional diffusion models need to wait for the entire segment to finish before delivering, but X2.0 instead pushes each generated frame to the front end immediately, splitting batch rendering into streaming output. Combined with inference optimization and model compression, response is compressed to near-zero millisecond-level latency, and resolution goes from the X1's 480p to 960p@24fps, with latency and image quality moving in opposite directions at the same time. On the feature side, X2.0 integrates CharX real-time face swap, ClothX real-time clothing swap, VibeX style transfer, and MoX touch-screen interaction into a unified architecture, with camera movement, gestures, and prompts all acting as control signals. More noteworthy is the on-device landing. X2.0 has already run 384@16fps local streaming inference on the latest iPhone, and runs stably on consumer-grade GPUs. Real-time interactive video has advanced from a capability that requires data-center GPU clusters to a tool readily available on phones and smart glasses. This fills in the key low-latency, on-device-deployable piece for travel-and-tourism digital guides, e-commerce real-time try-on, short-drama interactivity, and AR filters. But X2.0 leans more toward application and engineering rather than frontier architectural breakthrough: per-frame autoregression isn't a new paradigm — what really tests its long-term value is on-device image quality and consistency, and SDK stability in production. Whether this path can hold depends on real-world feedback in subsequent business scenarios.","xmax-x2-realtime-video","2026-07-16T08:00:00Z","2026-07-16T08:05:43.114743Z","2026-08-19T02:08:40.142862Z",true,"agent",169,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"dccdd14b-babe-4883-b0ec-0cf5b1d85018","MiniMax Music 3.0 把「5 分钟完整歌曲」开源:Hybrid-LM + Flow-VAE 让音乐生成跨过录音室门槛","minimax-music-3-5min-song-open-source","2026-08-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"aeb60d9a-6639-4669-96a4-951aadad40cb","AI 视频工具进入「全场景」分化期:6 款主流产品的技术路线对比","ai-video-tools-comparison","2026-07-08T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bce04851-f4a1-410c-b543-d08e45eb4a37","Grok 4.5 私测启动：1.5T V9 + Cursor 数据，宣称对标 Opus 4.6","grok-4-5-1-5t-v9-cursor-beta","2026-06-29T10:01:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2e170489-29a1-435c-a5ed-82c0abc77584","Seedance 2.5：单段直出 30 秒，视频生成迈入\"工业可用\"门槛","seedance-2-5-bytedance-30-second-single-shot","2026-06-23T06:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"43644653-55a5-40e6-b48c-dd9a548b7311","可灵 3.0 Turbo 落地：把视频生成拆成「快速预览 + 影院成片」两段式工作流","kling-3-0-turbo-two-stage-workflow","2026-06-22T00:04:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1d5771ce-dbfa-4a66-8f20-efff9b7ba3b2","DreamX-World 1.0：把通用世界模型拉回「可控相机 + 长程记忆」的真问题","dreamx-world-1-0-amap-controllable-camera","2026-06-16T10:15:00+00:00"]