Working agents have to read real files, coordinate tools, and ship real deliverables — but the training data behind them is mostly synthetic, and the verifiers are usually decoupled from the task itself. The 30 September paper from USTC, GraphForge, attacks that gap with one move: anchor both the task statement and each rubric criterion to the same evidence graph over real workspace files.
Evidence graph as task + verifier anchor
The pipeline runs in four steps. A seed is sampled from an occupation-grounded distribution to control task diversity; for each seed, GraphForge assembles a workspace of real files and builds an evidence graph over their relations; task requirements and rubrics are derived from that graph, so every criterion is naturally traceable to the files needed to verify it; an initial rollout tests executability and a revision agent repairs the task and rubrics against the original files before trajectories are collected. The paper lives at arXiv:2609.38923.
2,169 trajectories, three benchmarks move
The numbers are small-scale but hard. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories under OpenHands takes GDPVal to 1445.7 (+65.7). Under Claude Code, Workspace-Bench-Lite reaches 63.7 (+7.7) and SpreadsheetBench II reaches 24.0 (+13.7). A rejection fine-tuning pass on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks — suggesting the rubrics themselves are the useful artefact. This is consistent with the wave of working-agent papers (Terminal-Universe, SkillGym, CompoWorld) that all pivot the same way: away from environment scale, toward environment-verifier traceability.
Data and weights are open
The artefacts are fully open. The team has released a Hugging Face collection at groundhogLLM/graphforge containing the 2,169-trajectory SFT dataset (GraphForge-SFT-2169), a 27B SFT model (GraphForge-Qwen3.6-27B-SFT), and a 35B-A3B MoE variant SFT (GraphForge-Qwen3.6-35B-A3B-SFT). The HF papers page currently lists the work at rank #3 with 144 upvotes. Authorship spans the USTC system (Tao Gui, X.F. Zhao are recognisable names in LLM training) and groundhogLLM, the Hugging Face submitter (paper author Qisheng Su).
Why this matters for industry readers
The broader takeaway: as the chat-and-write capability race across GPT, Claude, and Gemini matures, the next real LLM battleground is whether models can actually do real work. The bottleneck is no longer model size — it is the realism of the task-verifier loop. GraphForge's evidence-graph approach is one specific answer; Terminal-Universe's reverse-engineered trajectories and SkillGym's auto-verifiable environments are different answers to the same problem. For anyone tracking agent infrastructure and training data, GraphForge is a clean example of that "rubric-traceable task synthesis" idea moving from a paper concept to a downloadable 27B checkpoint you can fine-tune on.