Video generation has been a crowded race this year, but a survey posted to arXiv on September 23 pulls the industry's attention back to a more fundamental question: a model can render beautifully, yet if it cannot remember what happened moments ago, long videos still fall apart.
Titled "The Past Frames the Future: Memory for Autoregressive Video Generation," the survey lists 25 authors. First author Harold Haodong Chen states on the Hugging Face paper page that this is the first systematic survey on memory mechanisms for AR video generation. Within two days the paper collected 35 upvotes on Hugging Face Daily Papers, and the companion awesome-list repository has already reached 54 stars — a clear signal of how much the community wants this problem named.
Why Memory Is the Choke Point of AR Video Generation
Autoregressive (AR) video generation extends visual sequences through causal rollouts: each step predicts the next segment based on what has been generated. The catch is that deployed models must operate under strictly bounded context windows, storage, and compute. As the survey puts it, three kinds of critical historical information — entity identities, dynamic states, and intervention-induced causal changes — often leave the active context long before their relevance diminishes.
The failure mode is familiar to anyone who has watched generated videos: a character turns around and comes back with a different face; a shattered glass un-shatters itself; the camera cuts away and returns to a room whose layout has changed. These issues used to be filed under "consistency." The survey reframes them all as a memory problem: persistent historical information maintained across outer AR steps that can still influence future generation even after the originating evidence is no longer locally accessible. That reframing is itself a contribution — it pulls scattered work on "long video," "world models," and "consistent generation" into one coordinate system.
Four Memory Carriers and a Taxonomy Map
The survey organizes the literature along five perspectives: Forms, Functions, Operations, Learning, and Evaluation. Judging from the companion repository's structure, the "carriers" of memory fall into at least four families:
- Visual Memory: retaining pixels or VAE latent representations directly, as in WorldMem, DecMem, and MemLearner
- Implicit State Memory: hiding inside attention caches, recurrent networks, or state-space model states
- Explicit State Memory: explicitly modeling entity-centric or spatial-geometric states
- Adaptive Parametric Memory: rewriting the model's parameters themselves
This is not an academic exercise in tidiness. Of the 282 entries in the companion repository, 181 are from 2026 — nearly two-thirds of the field's core work appeared in the past nine months alone. The competition between carrier routes has barely started, which is exactly when a map is most valuable.
The Real Shortfall: Benchmarks
The most sobering part of the survey is its verdict on evaluation: existing benchmarks cannot diagnose "genuine memory capability." The repository catalogs 14 memory-oriented benchmarks — R2M-Bench, MBench, WBench, MemoBench, and more — nearly all born in 2026, most covering only a single dimension. When vendors claim "minute-scale generation," what proves the character on screen still remembers the plot from three minutes ago? Without standardized evaluation, such claims rest mostly on self-certification.
So What
For teams building video generation, this survey and its companion repo (github.com/HaroldChen19/Awesome-AR-Video-Memory) belong in your bookmarks: check the four-carrier taxonomy before picking a route, and avoid re-investing in branches already shown to be dead ends. For observers, the more interesting signal is this — the "context window vs. external memory" debate that LLM circles argued over for two years is now replaying in video generation, at a much faster tempo. The question to watch: will one of these 14 benchmarks produce video's "MMLU moment," when memory capability becomes comparable for the first time?
References: paper arXiv:2609.28466 | repo Awesome-AR-Video-Memory