[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-onestreamer-4b-streaming-video-memory":3,"topics-all":35,"news-related-36c71e4e-e398-4583-8710-1732aefff06a":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"36c71e4e-e398-4583-8710-1732aefff06a","OneStreamer:4B 流式模型先记再答,八榜最佳","南京大学 MCG 牵头 11 家机构开源 OneStreamer-4B:基于 Qwen3-VL-4B 的流式视频模型,边看边把细节写成时间对齐的文本记忆,回答历史问题不再回看旧帧。作者报告其在对比的八项流式基准上综合分全部最高,状态 token 只监督 27.5% 即超稠密监督,上下文 token 省 93.1%。","先记再答,是流式视频模型最难的一课。帧每一秒都在进来,而细节在问题到来之前就滑出了上下文窗口——等用户终于问起\"孩子能去哪儿看书\",拍到的阅读角早就不在画面里。沿用离线视频 QA 的老办法只有两条路:要么把全部历史帧留在上下文里,token 和显存随时间线性膨胀;要么只留最近几帧,\"过去\"直接失忆。OneStreamer 给出第三条路:让模型边看边把值得记的东西写成文字。\n\n## 一次十一校协作,把记忆做成生成任务\n\nOneStreamer-4B 由南京大学 MCG 团队牵头,联合 PJLAB、京东、上海交大等共 11 家机构完成,通讯作者王利民。模型基于 Qwen3-VL-4B-Instruct 构建,代码以 Apache 2.0 协议开源,4B 权重、OneStreamer-1M 数据集(超 100 万条记录)与推理、评测代码全部放出。\n\n核心是把感知、记忆、响应统一进同一个生成过程:模型一边看,一边用两个控制符写笔记——`\u003C\u002FObserve>` 记局部细节,`\u003C\u002FSummary>` 总结已完成事件。笔记带时间戳留在文本历史里,等源帧滑出视觉窗口后依然可用;回答历史问题时靠这些记录补充最近视觉窗口,不重新访问历史视觉特征。\n\n## 只监督 27.5% 的状态 token,反而更强\n\n流式交互还有个隐蔽难题:模型大多数时刻应该沉默,重复的等待状态会在监督里占主导。PSTL(Proactive State Transition Learning)保留全部输出锚点,只挑代表性的状态变化与维持 token 监督——只用 27.5% 的标注状态 token 就超过稠密监督:ProactiveVQA、OmniMMI、OVO-Timing 三个基准拿到 48.7、36.6、41.6,全 token 稠密 CE 监督只有 26.1、30.8、1.5。\n\n## 记忆的价值,有数字为证\n\n官方项目页给出三组实测(单卡 H200):\n\n- **记忆消融**:同样只看最近 16 帧,加上 caption 记忆(PHCM)后,OVOBench Backward ASI 从 63.5 升到 71.6——超过全量历史的 67.6,实时感知分还从 80.9 升到 81.4,两项兼顾;\n- **token 效率**:360 秒样本上,上下文 token 从 62094 降到 4308,省 93.1%;显存从 25.18 GB 降到 9.98 GB,首 token 延迟从 4.560 秒降到 0.124 秒;\n- **更新速度**:360 秒 StreamingBench 片段平均每次更新 0.636 秒,低于 1 秒输入间隔,边看边记不掉队。\n\n在对比的八个基准(OVOBench、StreamingBench、OVBench、ODVBench、ProactiveVideoQA、OmniMMI、OVO-Timing、ViSpeak)上,作者报告 OneStreamer-4B 全部取得最高综合分。\n\n## 评论:记忆不是缓存,是写作\n\n这份工作最值得琢磨的转向,是把\"记什么\"从工程问题变成模型自己学的生成问题。文本可读只是副产品,真正的好处是文字天然带时间戳、可复用、便宜——4308 个 token 装下 360 秒视频里全部\"值得记的东西\",这个压缩比本身就是答案。两点保留:八榜最佳是作者对比集合内的结论,不是跨厂商公论;caption 记忆的上限取决于模型自己的判断力,写错的笔记比没有笔记更难纠正。对做实时视频 Agent 的团队,这条路线至少证明:与其无限扩视觉上下文,不如教会模型做笔记。\n\n参考:arXiv:2610.01762 \u002F mcg-nju.github.io\u002FOneStreamer \u002F github.com\u002FMCG-NJU\u002FOneStreamer","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01762","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"57ec78c8-35f1-48d1-8681-1a4652f15c1a","en","OneStreamer: 4B Streaming Model Remembers, Then Answers","NJU's OneStreamer-4B writes caption memory live while streaming video, answering without revisiting frames. Best on 8 benchmarks; 93.1% fewer tokens.","Remember first, answer later — the hardest lesson for streaming video models. Frames arrive every second, and details leave the context window before any question about them is known. When the user finally asks \"where can my kids go to read?\", the reading corner has long scrolled out of view. The two inherited options from offline video QA are both painful: keep every historical frame in context and watch tokens and memory balloon over time, or keep only the most recent frames and simply forget the past.\n\nOneStreamer offers a third path: let the model take its own notes while watching.\n\n## An eleven-institution effort that turns memory into generation\n\nOneStreamer-4B is led by the MCG group at Nanjing University together with PJLAB, JD, SJTU, USTC, CAS, CUHK, PKU, THU, FDU and ZJU — eleven institutions in total, with Limin Wang as corresponding author. Built on Qwen3-VL-4B-Instruct, the code is released under Apache 2.0, and the 4B model weights, the OneStreamer-1M dataset (over one million records), plus inference and evaluation code are all public.\n\nThe core idea unifies perception, memory, and response in one generation process. While watching, the model writes notes with two control tokens — `\u003C\u002FObserve>` for local details and `\u003C\u002FSummary>` for completed events. These time-grounded records stay in the text history after their source frames leave the visual window. When a question about the past arrives, the model relies on these text records to complement the recent visual window, without revisiting historical visual features.\n\n## Supervising 27.5% of state tokens beats dense supervision\n\nStreaming interaction hides another trap: the model should stay silent most of the time, so repeated waiting states dominate the supervision. OneStreamer's Proactive State Transition Learning (PSTL) keeps all output anchors but selects only representative state-change and state-persistence tokens. With just 27.5% of annotated state tokens supervised, it outperforms dense supervision: 48.7, 36.6 and 41.6 on ProactiveVQA, OmniMMI and OVO-Timing respectively, versus 26.1, 30.8 and 1.5 for all-token dense CE.\n\n## The numbers behind the memory\n\nThe project page reports three sets of measurements (single H200):\n\n- **Memory ablation**: with the same recent-16-frame window, adding caption memory (PHCM) lifts OVOBench Backward ASI from 63.5 to 71.6 — above even the full-history setting at 67.6 — while the real-time score rises from 80.9 to 81.4. Memory does not tax real-time perception; it improves both at once.\n- **Token efficiency**: on a 360-second OVOBench sample, context tokens drop from 62,094 (full history) to 4,308 — 93.1% fewer; GPU memory falls from 25.18 GB to 9.98 GB, and time-to-first-token from 4.560 s to 0.124 s.\n- **Update speed**: on a 360-second StreamingBench clip, PHCM averages 0.636 s per update, below the one-second input interval — it keeps up while watching.\n\nAcross the eight streaming video benchmarks it compares against (OVOBench, StreamingBench, OVBench, ODVBench, ProactiveVideoQA, OmniMMI, OVO-Timing, ViSpeak), the authors report the best aggregate score for OneStreamer-4B on every one.\n\n## Commentary: memory is not caching, it is writing\n\nThe most thought-provoking shift here is turning \"what to remember\" from an engineering question — how many frames to keep, how much to compress — into a generation task the model learns itself. Legible text is a side benefit; the real win is that text is naturally time-stamped, reusable, and cheap. Fitting everything \"worth remembering\" from 360 seconds of video into 4,308 tokens is itself the answer. Two caveats: \"best on eight benchmarks\" is a claim within the authors' own comparison set, not a cross-vendor verdict; and caption memory is bounded by the model's own judgment — a wrong note is harder to fix than a missing one. For teams building real-time video agents, the lesson stands: rather than endlessly extending visual context, teach the model to take notes.\n\nReferences: arXiv:2610.01762 \u002F mcg-nju.github.io\u002FOneStreamer \u002F github.com\u002FMCG-NJU\u002FOneStreamer","onestreamer-4b-streaming-video-memory","2026-10-02T15:07:48Z","2026-10-02T15:08:49.711153Z","2026-10-02T15:08:49.711165Z",true,"agent",32,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"bbe8d55a-1069-42fa-a342-d944d53fdb4b","操作电脑的 27B 开源权重模型 Holo4:最强版禁商用","holo4-open-weight-computer-use","2026-10-02T13:12:01+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"dc61debd-1d77-4d5a-9d59-5b23c3da07de","蚂蚁开源Realtime-Venus：9B全双工模型边说边干活，三项续聊指标超GPT-4o","ant-realtime-venus-full-duplex-delegation","2026-09-30T23:10:53+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"0d357a0b-42da-40af-8038-e8c035cb9810","Apple 开源 LensVLM-9B:先扫压缩图,再读原页","apple-lensvlm-9b-weights-huggingface","2026-09-26T15:20:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00"]