A 47-page paper from Amazon landed on arXiv on September 24, 2026, and it does something rare: it documents a complete post-training recipe on an open base model, in full. The team built Rufus-Air on top of GLM-4.5-Air-Base, the open Mixture-of-Experts checkpoint with 106B total and 12B active parameters released by Zhipu, and published the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the whole run. Twenty-two authors, all at Amazon, listed alphabetically by surname. In an era when most labs treat post-training as a trade secret, this is an unusually transparent artifact.

The skeleton: one serial pipeline, eight stages

The recipe is a single serial pipeline with no branching: SFT, then Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and finally RLHF. Each stage trains from the checkpoint the previous one produced. Two axes set the order. On capability, stages move from basic to advanced, so each builds on what the previous one seeded. On reward type, stages run from hard, verifiable rewards toward softer judge-based signals, which limits how long the policy is exposed to reward hacking.

The SFT numbers are concrete: 9.01M samples, 44.5B raw tokens, 66.7M conversational turns, and 27.0B supervised tokens after masking system messages, user turns, and tool observations. Decontamination used a word-level 8-gram overlap screen and removed 3,529 samples in total, 3,321 of them from a single synthetic text-to-terminal corpus. The SFT checkpoint alone (step 3799) already leads the public GLM-4.5-Air release by 5.3 points on IFEval and 24.2 on IFBench. The authors push back explicitly on treating SFT as a warm-up: it establishes the capability floor that RL later refines.

Stagewise gains, measured honestly

Difficulty filtering is one of the recipe's practical workhorses: prompts the policy already solves at a rate above 0.8 are dropped as too easy, and prompts with zero observed success are dropped as currently unlearnable, so training concentrates on the learnable band in between. Reasoning RL uses GSPO as the policy-gradient backbone and lifts GPQA from 68.2 to 73.5 (+5.3), at the cost of 2.6-2.8 points on both AIME years — a stated trade-off, not an accident. Coding RL raises LiveCodeBench v6 from 67.9 to 74.5, peaking at 75.9 under an extended 128K response budget. Instruction-Following RL trains with GRPO and moves IFEval up 4.0 points to 94.5 while leaving GPQA (+3.2) and AIME 25 (-1.2) essentially level.

The three agent stages carry the most new information. General Agent trains on 10K single-server MCP tasks across 1K synthetic environments, and the authors report MCP-Atlas +7.80 and Tau2-Retail +9.80, reading this as evidence that synthetic MCP environments provide a general prior for tool orchestration. Coding Agent runs Docker-sandboxed tasks with test-suite rewards on the Harbor task format, moving SWE-bench Verified from 65.60 to 67.80 and Terminal-Bench 2.1 from 38.76 to 40.17 — with an honest caveat that the stage was trained on the compute available rather than to convergence. Search Agent works against the open web, iterating up to 100 tool calls within a 128K-token budget during training, and adds +3.0 on BrowseComp, +5.4 on Seal-0, and +3.4 on HLE-Verified. The final RLHF stage uses the open Skywork-Reward-V2-Qwen3-8B reward model and only the prompts from HH-RLHF, lifting Arena-Hard v2 Hard Prompt from 83.06 to 89.05 and Creative Writing from 38.56 to 52.97.

How the finished model stacks up

Under the paper's own harness, with four models evaluated under identical conditions, Rufus-Air leads the official GLM-4.5-Air release on every reported benchmark except Arena-Hard v2 Creative Writing: IFEval 95.4 vs 83.0, SWE-bench Verified 65.6 vs 50.6, Arena-Hard v2 Hard Prompt 89.1 vs 55.0. Same-base INTELLECT-3, plus Nemotron-3, GPT-OSS, and Qwen3.5 at similar scale, appear in the same table; on competition math and knowledge, the authors' own reading is that the recipe moves most where it spends the most training signal. One caveat the paper flags only lightly: the official GLM-4.5-Air reference dates from July 2025 and predates several of the newer benchmark generations, so the comparison's headroom partly reflects benchmark drift.

Why this recipe matters

Three layers. First, the information asymmetry around post-training is being filled in. Open recipes such as Tülu and OpenRAID exist, but a fully documented run spanning MoE architecture, agent environment synthesis, and budget-conscious RLHF on a single open base remains scarce; Amazon has effectively standardized the engineering path from open base to finished model once. Second, it is a subtle signal for the Chinese open ecosystem: a US hyperscaler chose Zhipu's GLM-4.5-Air-Base as its training ground, and the result outperforms the official release on multiple benchmarks — both the base value and the post-trainability of open-weight models got validated at once. Third, openness itself may become a competitive axis: as benchmark scores converge, a reproducible recipe with negative results spelled out is worth more long-term than another leaderboard entry.

The so-what for ordinary developers: RL stages that fit on 8-32 nodes mean mid-sized teams, not just frontier labs, now have a path to copy. The sandbox service runs on the order of ten thousand concurrent sandboxes at roughly $10K per month, and search API calls are billed per use — budget these separately. The recipe paper is not the endpoint; it is one step in moving post-training from alchemy toward engineering. The next thing to watch is how quickly the community reproduces a second batch of models from it.

References: full paper at arXiv:2609.29421 (https://arxiv.org/abs/2609.29421); Hugging Face paper page at https://huggingface.co/papers/2609.29421 .