Long-horizon agents keep tripping over their own history: unverified assumptions and outdated plans persist in context and distort later decisions. A Renmin University of China team names this task-state contamination and proposes AEWM, which reframes the language world model objective from predicting environment observations to editing task state. Action Judge classifies decisions as Critical, Exploratory, or Noisy; State Revision rewrites noisy continuations; EditAct wires both into real execution. Self-reported: 70.5% macro-F1, 10.6 points above the strongest frontier baseline, and 3.2-6.7 point average gains across six benchmarks and three agent backbones.

A world model that stops predicting the environment

The paper opens by questioning the mainstream approach: existing language world models learn to predict the next environment observation, yet tool responses are high-entropy and execution-dependent, so reconstructing them adds limited value when real feedback is cheap. AEWM instead models how an agent's reasoning and actions shape future task progress.

The design has two parts. Action Judge classifies decisions before execution into Critical, Exploratory, and Noisy categories. State Revision rewrites noisy reasoning-action continuations from the same observed history. The inference framework EditAct integrates both with real execution: rather than standing aside and offering critiques, it directly changes the state that subsequent decisions depend on. A further variant, AEWM-RFT, fine-tunes on verified EditAct trajectories via rejection sampling, beating Self-RFT by 2.2-2.6 points across three domains without online AEWM guidance.

Self-reported numbers: +10.6 on the judge, +3.2-6.7 across six benchmarks

The team trained AEWM across Search, Terminal, and Software Engineering through mid-training plus supervised fine-tuning. On their own Action Judge benchmark, AEWM reaches 70.5% macro-F1, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2 to 6.7 points over the strongest baseline.

One caveat: all figures come from the authors' own evaluation; no independent third-party replication is out yet, so treat them as self-reported.

Why this matters for agent engineering

The core value is moving the correction point earlier. The traditional loop lets errors happen, then retries or reflects; AEWM identifies and rewrites noisy decisions before they contaminate the history. As context windows stretch and task chains grow longer, history hygiene may prove cheaper than larger models: contamination accumulates with steps, while parameter count buys no immunity.

A signal worth noting: the author list comes from Renmin University's RUC team (corresponding authors Wayne Xin Zhao and Ji-Rong Wen, among others), and academic output on agent infrastructure is visibly concentrating in Chinese teams.

Full paper at arXiv:2609.28416. The takeaway: next time your agent cites a step-3 wrong assumption at step 20, don't rush to swap in a bigger model. It may just need a state edit.