Multimodal post-training usually treats distillation and reinforcement learning as two sequential stages: distill first, then RL. The Pistis technical report from a ByteDance team proposes a different answer—IDRL (Interleaved Distillation and Reinforcement Learning) puts both objectives into a single training loop and alternates between them, instead of chaining the stages or merging them into one static weighted loss.
The models first
The Pistis family covers two multimodal scales, 27B and 9B, built on Qwen3.6-27B and Qwen3.5-9B respectively. Each scale ships two variants: Pistis-Thinking targets deep multimodal reasoning, while Pistis-Agentic additionally consumes agentic trajectory data for long-horizon planning, iterative reasoning, and tool use. The SFT stage uses roughly 3.2M multimodal QA pairs, with agentic trajectories split into about 40% tool-integrated reasoning, 20% search, and 40% general agent trajectories.
On 24 non-grounding benchmarks, Pistis-27B-Thinking is comparable to Qwen3.8-27B (82.3 vs 82.4), while achieving the highest grounding average of 80.5, exceeding Qwen3.8-27B by 8.6 points. The Agentic variants are particularly strong in multimodal search: relative to Qwen3.6-27B, Pistis-27B-Agentic improves BrowseComp-VL, MMSearch, VDR-testmini, and LiveVQA by 8.8, 3.3, 3.2, and 9.7 points respectively.
IDRL: alternation, not weighting
The core idea in one sentence: RL sharpens the policy toward high-reward outputs and reduces entropy, while on-policy distillation pulls the student back toward the teacher's distribution and preserves entropy. Summing the two objectives in the same step makes their gradients fight—the report derives that when the student is already more confident than the teacher on a token, the two contributions point in opposite directions, and the summed update can even lower the reward objective. IDRL instead alternates by cycle, for example 5 OPD steps and 5 RL steps per cycle, optimizing exactly one objective per step; once entropy plateaus, it switches to pure RL. The 9B variants use IDRL while the 27B variants are trained with pure RL—and double as frozen OPD teachers: 27B-Thinking teaches 9B-Thinking, and 27B-Agentic teaches 9B-Agentic.
For long-horizon agentic tasks there is a dedicated design: PAS (step-level positive-advantage suppression). Trajectory-level rewards only see final success, which lets lucky-but-useless intermediate actions earn positive credit. Pistis zeroes the positive advantages of rejected, repeated, or post-answer tool-call steps, while keeping penalties in failed trajectories—cleaner credit assignment.
PAH: optimize the harness, not the parameters
The report also has a system-level bonus: PAH (Pistis-Auto-Harnessing). The model and tool interface stay frozen; the optimization target is the surrounding inference harness. An Optimization Agent runs a five-stage closed loop on a development set: attribute failures, propose one falsifiable change, implement it, gate it with a small canary run, then evaluate on the full development set—rolling back whenever the development metric does not improve. The resulting harness ships a structured Candidate Ledger, conditionally loaded Search Skills, evidence-driven checkpoints, and budget-aware convergence. It is validated on VDR-testmini under the same interaction budget, improving performance without extra model updates or added interaction budget.
So what
The value of IDRL is not the "distillation plus RL" combination itself, but how the two are combined: alternating rather than weighting is essentially about separating two opposing gradient objectives in time, making entropy a renewable exploration resource. For post-training teams, this is a reusable scheduling recipe; for observers, the signal worth remembering is that post-training competition has shifted from "what data to use" toward "how to orchestrate training signals"—and the harness itself is becoming an optimization target. (arXiv:2609.28554)