In the world of ExplorationBench, EMIT(100) does not print 100 — it prints 127, because integer literals in this alien language are silently XOR-ed with 27, and PLUCK counts positions from one although the manual says zero. The manual is deliberately wrong. The only way through is to experiment.

That is the design intent of the new benchmark. A 20-author research team, with Tencent Hunyuan among the signing institutions, turned the fuzzy question of whether AI systems can genuinely explore into an executable, exactly verifiable measurement framework. The paper, ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds, hit arXiv on Sep 24 (2609.30199) and has since entered the Hugging Face daily papers list.

Why alien worlds

Evaluating exploration has a built-in deadlock. Tasks must be new to the model, otherwise you are testing recall; yet answers must be exactly verifiable, otherwise you cannot grade. Math and coding satisfy the second requirement but not the first — pre-training data may already contain the answers. Genuine scientific discoveries satisfy the first but not the second — verification can take expert years.

The team's answer is to build two alien worlds. AlienCode is a small calculation language with 31 hidden discovery targets and 70 held-out tasks; AlienLogic is a natural-deduction system with patched inference rules, 24 discovery targets, and 70 tasks. Both are deterministic and executable: an interpreter grades programs, a proof-checker grades proofs, and no LLM judge is involved. Each system receives a flawed manual and a few worked examples, then explores for four rounds — submitting programs or proofs, reading results, revising hypotheses. After each round it is tested closed-book, with tools disabled.

Results: exploration works, unreliably

Ten frontier systems, three independent trajectories each, produced a dense set of findings:

  • Before exploration, no AlienCode trajectory exceeds 15.7%. After four autonomous rounds, Best@3 peaks at 87.6% (Claude Opus 5), and 7 of 10 systems pass 60%.
  • Remove environment feedback and let models deliberate for the same number of turns: scores collapse to 0.5%-11.0%. More thinking is not more knowing.
  • Who designs the experiments matters. Median accuracy is 66.0% under autonomous exploration, 40.7% when a system's own best probes are replayed to it, and 5.7% under a fixed probe sequence. Identical evidence — the gap comes from designing experiments itself.
  • Rankings do not transfer. Grok 4.6 is fifth in AlienCode but first in AlienLogic; DeepSeek-V4-Pro bottoms out at 12.9% in AlienCode yet reaches 73.8% in AlienLogic. The Spearman correlation between the two boards is 0.35 — exploration ability is not one portable score.

Three negative findings cut deeper. Saying is not using: on tasks whose required rules a system states correctly, it still solves only 70.9%. Exploration is unstable: trajectories of one system under one budget end up to 72.8 points apart, and 6 of 30 AlienCode trajectories finish at least 3 points below an earlier milestone. And in AlienLogic, simply handing over the complete rule set yields 93%-97%, still beating every system's autonomous exploration.

So what

For the AI-scientist narrative, this is a temperate bucket of cold water. Frontier models can genuinely reverse-engineer a set of unfamiliar rules from scratch within a bounded budget — that part is real. But rerun the same system and the result may shrink sharply; reciting a discovered rule does not guarantee applying it. Treating this benchmark as a measurement of exploration being unreliable and non-transferable is probably more valuable than farming it as another leaderboard.

Reference: arxiv.org/abs/2609.30199