Long-horizon agents keep making the same move: picking a path at a fork. Which hypothesis to test, which implementation to build on, which fix to attempt — these mid-run decisions decide the whole trajectory. A Microsoft team calls this ability to pick the better direction before the outcome shows "taste," and they just released a benchmark that measures it: Taste-Bench, in the paper "The Tasteful Agent" (arxiv.org/abs/2609.25804), now #1 on Hugging Face Daily Papers with 61 upvotes.
What taste means here
Each question is deliberately spare: the model gets the task, the full trajectory up to a decision fork, and two candidate next steps. The hidden rest of the trajectory proves which one is right. A wrong choice often looks reasonable in the moment and burns most of the budget later. The authors evaluated 14 frontier models; the best, GPT-5.6 Sol, answers only 59.7% correctly, against 25% for random guessing. Flagship judgment at real forks is nowhere near reliable.
Two counterintuitive findings stand out. First, the later the deciding evidence appears in the trajectory, the worse models do: accuracy falls from 62.3% to 21.0%. Second, a larger reasoning budget does not rescue it — more thinking tokens do not improve choice accuracy. Taste is not something a model earns by thinking longer.
How the 502 questions were built
The benchmark needs no human annotation. The authors mined forks from SWE-bench, SWE-bench Pro, and METR's MALT release of RE-Bench and HCAST trajectories: parallel attempts at the same task that diverge at the same point with different recorded outcomes, and detours inside a single run where the agent abandons a direction after an observed failure and recovers. The later part of a trajectory labels the earlier decision for free.
Filtering is strict: a question is dropped as trivial when every judge model answers it from the candidate wording alone, and as undecidable when any judge disagrees with its label after reading the full record. Of 4,657 mined forks, 502 survive — 390 engineering and 112 research questions. The protocol also blocks position bias: every question is asked in the published order and in its exact reverse, and only double-correct answers count; a model that always picks the same position scores 0. The authors report 98.8% agreement with human review.
Taste is trainable
The leaderboard itself is informative: GPT-5.6 Sol (59.7) and GPT-5.5 (59.5) lead, Claude Opus 5 sits at 55.5, GLM-5.2 at 53.9, DeepSeek V4 Flash at 43.3, and Grok 4.20 Reasoning bottoms out at 15.7 — 459 of its 1,004 responses were unparseable and count as wrong.
The real hook: taste can be trained. The authors distilled the judgment of a teacher that had seen the outcome into a student; Qwen3.6-27B goes from 30.0% to 47.9% on unseen tasks. Let that student advise an agent, and end-to-end success on SWE-bench Pro jumps from 14.6% to 33.7%. Judgment can be distilled from hindsight and fed back into running agents.
So what
For two years the field has optimized whether agents finish the course, and benchmarks have watched end-to-end success. Taste-Bench points the camera at every mid-run turn, and the data says that is exactly today's thinnest weak spot — the best model is under 60%, and brute-force reasoning budgets do not help. For agent engineers, pausing at a fork to ask a hindsight-distilled advisor may beat swapping in a bigger model. Code is MIT, data CC BY 4.0, with a gated dataset release to limit training contamination; a full run is 1,004 requests and about 8M input tokens. Worth running against your own model to see how much taste it has.