Human top players score a peak of 84.3 on unfamiliar text games; Claude Opus 5.5 reaches 80.1 — a 4.2-point gap. Averaged over whole runs, though, the picture flips: Opus 5.5 averages 61.0, higher than the human Top-1 reference at 53.5. These numbers come from Learn2Play Bench (arXiv:2610.08215), released October 8 by Bryan Hooi's group at the National University of Singapore; it climbed to #2 on Hugging Face Daily Papers with 131 upvotes.

Closing off the lookup path

Existing agent benchmarks share a flaw: task rules are either spelled out in instructions or already seen in pretraining data, so you cannot tell learning from recall. Learn2Play ships 20 newly designed text-game templates whose rules are deliberately novel or counterintuitive and can only be discovered through trial and error — model weights stay frozen, and experience lives entirely in context. Seeds generate game instances, feedback is reproducible, scoring is automatic, and some games reshuffle the visible situation between attempts while keeping the rules fixed, specifically testing transfer. The protocol is explicit: one selected rule seed per template, 3 independent trials, 5 episodes for normal games and 10 for challenge games, scores normalized to a reference ceiling, with challenge games weighted 3 against 1.

The leaderboard: average flipped, peak still short

Across the 15-row table (11 OpenCode backbones plus 4 human references), three rows stand out. Claude Opus 5.5 peaks at 80.1 and averages 61.0 — that average beats every human reference, including Top-1, which patches together each game's highest-scoring player (average 53.5). Human Top-1 still holds the peak at 84.3. The steepest learner in the backbone table is Qwen 3.8 Max with a learning gain of +25.2. GPT-6 Astra ranks second among models with a 76.2 peak; Kimi K3 sits at 68.5, DeepSeek V4 Pro at 54.2, and GPT-5 Mini at the bottom with 34.0 — below even the human-mean reference of 43.6.

Fix the model, swap the harness, gain 7 points

The third finding matters most to practitioners: with the backbone fixed, changing the agent harness can improve performance while cutting estimated inference cost. The project page's cost-performance chart cites the extreme case — Claude Code with Opus 5 peaks at 81.8%, more than 7 points above Opus 5 running bare as a backbone at 74.5, and just 2.5 points from human Top-1. Astra's bill is transparent too: 150 US dollars in total across 600 verified completed episodes, or 0.25 dollars per episode.

Full logs beat clever summaries

Two counterintuitive observations deserve their own paragraph. First, feeding complete records of actions and feedback back into context supports better learning than first summarizing the experience into rules or strategies — the act of summarizing itself loses information. Second, human players hold the higher peak while exploring more varied strategies and repeating fewer actions; agents fall into the same pits over and over.

Three buckets of cold water before reading

The project page annotates itself: the leaderboard is a manually curated snapshot (updated October 8), not auto-refreshed; all numbers are aggregate point estimates without significance testing; the human references are per-game strongest-player patches, not average people, and humans played without AI tools or code. The code repository is currently an anonymous-review snapshot — 12 stars, 3 commits — and the 20 game templates run on Python 3.10+ with the standard library alone (for example, python play.py poisoner --seed 2); the play portal learn2play.fun requires sign-in.

References: paper and leaderboard at arxiv.org/abs/2610.08215, project page liushiliushi.github.io/learn2play-bench-website, code at github.com/liushiliushi/Learn2Play-Bench.

Matching or beating the human averages means models have gotten decent at learning stably. The missing 4.2 points of peak score test whether, on first contact with a counterintuitive rule, an agent dares to try more paths — so far, that remains the human home ground.