Remember first, answer later — the hardest lesson for streaming video models. Frames arrive every second, and details leave the context window before any question about them is known. When the user finally asks "where can my kids go to read?", the reading corner has long scrolled out of view. The two inherited options from offline video QA are both painful: keep every historical frame in context and watch tokens and memory balloon over time, or keep only the most recent frames and simply forget the past.

OneStreamer offers a third path: let the model take its own notes while watching.

An eleven-institution effort that turns memory into generation

OneStreamer-4B is led by the MCG group at Nanjing University together with PJLAB, JD, SJTU, USTC, CAS, CUHK, PKU, THU, FDU and ZJU — eleven institutions in total, with Limin Wang as corresponding author. Built on Qwen3-VL-4B-Instruct, the code is released under Apache 2.0, and the 4B model weights, the OneStreamer-1M dataset (over one million records), plus inference and evaluation code are all public.

The core idea unifies perception, memory, and response in one generation process. While watching, the model writes notes with two control tokens — </Observe> for local details and </Summary> for completed events. These time-grounded records stay in the text history after their source frames leave the visual window. When a question about the past arrives, the model relies on these text records to complement the recent visual window, without revisiting historical visual features.

Supervising 27.5% of state tokens beats dense supervision

Streaming interaction hides another trap: the model should stay silent most of the time, so repeated waiting states dominate the supervision. OneStreamer's Proactive State Transition Learning (PSTL) keeps all output anchors but selects only representative state-change and state-persistence tokens. With just 27.5% of annotated state tokens supervised, it outperforms dense supervision: 48.7, 36.6 and 41.6 on ProactiveVQA, OmniMMI and OVO-Timing respectively, versus 26.1, 30.8 and 1.5 for all-token dense CE.

The numbers behind the memory

The project page reports three sets of measurements (single H200):

  • Memory ablation: with the same recent-16-frame window, adding caption memory (PHCM) lifts OVOBench Backward ASI from 63.5 to 71.6 — above even the full-history setting at 67.6 — while the real-time score rises from 80.9 to 81.4. Memory does not tax real-time perception; it improves both at once.
  • Token efficiency: on a 360-second OVOBench sample, context tokens drop from 62,094 (full history) to 4,308 — 93.1% fewer; GPU memory falls from 25.18 GB to 9.98 GB, and time-to-first-token from 4.560 s to 0.124 s.
  • Update speed: on a 360-second StreamingBench clip, PHCM averages 0.636 s per update, below the one-second input interval — it keeps up while watching.

Across the eight streaming video benchmarks it compares against (OVOBench, StreamingBench, OVBench, ODVBench, ProactiveVideoQA, OmniMMI, OVO-Timing, ViSpeak), the authors report the best aggregate score for OneStreamer-4B on every one.

Commentary: memory is not caching, it is writing

The most thought-provoking shift here is turning "what to remember" from an engineering question — how many frames to keep, how much to compress — into a generation task the model learns itself. Legible text is a side benefit; the real win is that text is naturally time-stamped, reusable, and cheap. Fitting everything "worth remembering" from 360 seconds of video into 4,308 tokens is itself the answer. Two caveats: "best on eight benchmarks" is a claim within the authors' own comparison set, not a cross-vendor verdict; and caption memory is bounded by the model's own judgment — a wrong note is harder to fix than a missing one. For teams building real-time video agents, the lesson stands: rather than endlessly extending visual context, teach the model to take notes.

References: arXiv:2610.01762 / mcg-nju.github.io/OneStreamer / github.com/MCG-NJU/OneStreamer