While most teams stockpile successful rollouts for distillation and throw failures away, a paper posted on arXiv on September 30 (2609.40111) argues the opposite: the Agent Error Dataset (AED) turns 50,228 error-diagnosis pairs into a reusable training asset. The opening line makes the thesis explicit — an unsuccessful rollout carries far more information than its final reward: the observations the agent saw, the actions it took, and the environment's responses are all reviewable evidence.

What is inside

The scale numbers are solid: 50,228 error-diagnosis pairs drawn from 9,961 source tasks, spanning 33 environments, 19 harness families and 23 policy models in text-based agent systems. After deduplication the collection holds 38,278 distinct source traces, averaging 1.31 diagnoses per trace — the same failure can be re-diagnosed from new angles without re-running the original rollout. Full execution metadata is retained for cross-setting failure analysis.

The five-stage pipeline

The companion AET (Agentic Error-to-Training) pipeline works in five steps. First it collects naturally occurring failures while screening out infrastructure faults and grader bugs, keeping only genuine agent errors. Second, each error gets a diagnosis: the failing step and the responsible agent, with an explanation that must cite the trace. Third, diagnoses pass structural and semantic grounding checks; rejected proposals are kept on record. Fourth comes replay — from the same checkpoint, under matched policy, harness, budget and verifier settings, the proposed correction races against a fresh original-action retry. Fifth, the data is organized into training views for diagnosis SFT, recovery SFT and action preferences.

The numbers

Across 3,062 matched replay pairs, first-proposal corrections lift verifier pass rates from 18.4% to 51.1% — a 32.7-point gain, roughly 2.8x for the same checkpoint. Fine-tuning Qwen3-8B on a frozen diagnosis release raises exact-step agreement with internal teacher labels from 47.2% to 63.6%, beating the strongest prompted reference at 54.7%. On the acting side, action-only repair training scores 6.67 points higher than success-only training on WebShop-lite (single seed), and diagnosis-guided continuation reaches 31.3% versus 26.0% for generic reconsideration and 18.6% for plain replay.

Caveats worth noting

The code, checkpoints and data are not out yet — the paper states they will be open-sourced after acceptance, subject to licenses and privacy review. The 6.67-point actor result is explicitly labeled single-seed. Corresponding author Heng Ji is affiliated with Apodex, and the paper acknowledges Tianqiao Chen for guidance and support — worth watching where this industrial-academic line goes next.

The industry signal is direct: while compute pours into synthetic success data, failure traces may be the free ore everyone is throwing away. A 2.8x pass-rate jump from corrected actions shows that knowing how to admit mistakes is itself a trainable capability. One question to leave you with: after your agents crash, where do those traces go now?

Paper: https://arxiv.org/abs/2609.40111