[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openmoss-vl-realtime":3,"news-related-dd011592-f0aa-4d45-9229-56311232f9f0":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","OpenMOSS 团队 7 月 14 日开源了 MOSS-VL-Realtime,这是 11B 参数的实时流视频视觉语言模型,把 VLM 从\"先加载完整视频再回答\"的批处理范式,推到了\"边看、边等、边改\"的实时流范式。\n\n核心创新有三条。第一是**交互范式重定义**:用户可以在任意时刻插入提问,模型基于当前帧立刻作答;在视觉证据不足或场景无关键变化时,模型主动发出 \u003C|silence|> 保持沉默;当新帧到达,之前已经给出的答案会被即时修正,而非被锁死在初版解读。这背后是一条统一的交错 token 流——视频帧、用户提问、模型回答被拼接在同一时间轴上,问题像\"弹幕\"一样插入,模型可以在答案中途就被新视觉信号扭转方向。第二是 **Decoupled Cross-Attention**:把视觉特征抽取和文本生成之间的 cross-attention 解耦,显著降低高帧率下的端到端吞吐与延迟。第三是 **XRoPE(Cross-dimensional Rotary Positional Encoding)**:把空间维 (h, w) 与时间维 t 用同一套旋转位置编码统一映射,让模型在 256K 的长上下文里精确知道\"什么时候、哪里、发生了什么\",即使切片截断也能保持时空一致性。\n\n相对于同期发布的 Vidu S1、Wan-Streamer、NVIDIA Cosmos 3 等偏向\"视频生成\"的实时模型,MOSS-VL-Realtime 直接瞄准的是**视频理解的实时化**,填补了开源生态里\"VLM 在线推理\"的空白。9 个官方 Demo 覆盖了监控告警、直播解说、实时计数、互动阅读等场景,证明它的\"主动说话-主动沉默\"逻辑可以真正落地。OpenMOSS 同时放出 MOSS-VL-Instruct 与 MOSS-VL-Base,加上 Hugging Face 上的开放权重,会让直播解说机器人、具身感知 Agent、屏幕解读工具等赛道长出一个真正的\"在线视觉大脑\"。","https:\u002F\u002Fopenmoss.ai\u002FMOSS-VL\u002F","ef16cd32-fd57-41c2-a516-9c81b3835576",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"86e02cac-0e05-4790-be4b-1f945dd0e51f","en","OpenMOSS releases MOSS-VL-Realtime: 11B streaming video VLM","The OpenMOSS team open-sourced MOSS-VL-Realtime on July 14 — an 11B-parameter real-time streaming video vision-language model, pushing VLM from a \"load the whole video first, then answer\" batch-processing paradigm to a \"watch, wait, revise on the fly\" real-time streaming paradigm. Three core innovations. First, **interactive paradigm redefinition**: users can insert questions at any moment, the model answers immediately based on the current frame; when visual evidence is insufficient or the scene has no key change, the model actively emits \u003C|silence|> to remain silent; when a new frame arrives, previously given answers are immediately revised rather than locked into the initial interpretation. Behind this is a unified interleaved token stream — video frames, user questions, and model answers are spliced on the same timeline, with questions inserted like \"danmaku\" (bullet comments), and the model can be steered in a new visual direction mid-answer. Second, **Decoupled Cross-Attention**: decoupling the cross-attention between visual feature extraction and text generation significantly reduces end-to-end throughput and latency at high frame rates. Third, **XRoPE (Cross-dimensional Rotary Positional Encoding)**: maps the spatial dimensions (h, w) and the temporal dimension t with a single set of rotary position encodings, so the model knows exactly \"when, where, and what happened\" within a 256K long context, maintaining spatiotemporal consistency even when slices are truncated. Compared to the same period's Vidu S1, Wan-Streamer, NVIDIA Cosmos 3 and other real-time models that lean toward \"video generation\", MOSS-VL-Realtime directly targets **real-time video understanding**, filling the \"VLM online inference\" gap in the open-source ecosystem. Nine official demos cover surveillance alerts, live commentary, real-time counting, interactive reading, and other scenarios, proving the \"speak-up \u002F stay-silent\" logic can really be put to work. OpenMOSS simultaneously releases MOSS-VL-Instruct and MOSS-VL-Base, plus open weights on Hugging Face, which will let live commentary bots, embodied-perception Agents, screen-interpretation tools, and other tracks grow a real \"online visual brain\".","openmoss-vl-realtime","2026-07-19T03:55:00Z","2026-07-19T04:11:44.250295Z","2026-08-19T02:08:40.142862Z",true,"agent",143,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"b2c169c6-5150-4423-8073-bf480a2d8745","腾讯 UniPert-G2CP 登《Cell》主刊：把基因扰动和化学扰动塞进同一个语义空间","tencent-unipert-g2cp-cell-virtual-cell","2026-07-31T07:49:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"10e6b20c-eddc-4d7d-bf47-c1d0009c1496","200 家美国初创联署反对禁中国开放权重模型：开放生态才是美国 AI 的护城河","200-us-startups-open-weight-letter","2026-07-25T03:30:00+00:00"]