How much of a high SWE-bench score comes from real reasoning, and how much from simply recognizing the repository? A paper submitted to arXiv on Aug 21 (arXiv:2609.27891) by a Shanghai Jiao Tong University-led team offers an experiment that separates the two: keep the tasks, but swap the face of the test repository, and see how much the agent still scores.

A repository that collapses only at evaluation time

The paper proposes SchrodingerRepo, whose core idea borrows from Schrödinger's cat: the test repository is treated as an evaluation-time latent variable, dynamically instantiated only when the agent enters the evaluation environment. The instantiated version preserves the original executable behavior while progressively eroding familiar cues through four transformation levels — problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting. Authors come from SJTU (co-first authors Silin Chen and Yufei Yang, corresponding author Xiaodong Gu) and Xi'an Jiaotong University; the code is open-sourced under MIT.

A motivation experiment first quantifies the leakage: for every evaluated model, more than 65% of SWE-bench Verified instances show clear data-leakage evidence, and more than 18% can be recalled at the patch or test level. This echoes OpenAI's earlier judgment that SWE-bench Verified no longer reliably measures frontier coding capability.

Four models drop together

Using the mini-swe-agent scaffold on SWE-bench Verified, the team evaluated four backends: GPT-5.4-mini, GPT-5.1, Gemini-3.1-Flash-Lite, and DeepSeek-v4-Flash. With all four transformation levels enabled, Pass@1 fell across the board: GPT-5.4-mini from 46.8% to 35.6%, DeepSeek-v4-Flash from 72.8% to 66.8%, GPT-5.1 from 44.6% to 36.2%, and Gemini-3.1-Flash-Lite from 56.7% to 42.3% — an overall drop of 6.0-14.4 percentage points, statistically significant (p<0.05).

The harshest level was not code rewriting but namespace remapping: that single level cut three models by 7.4, 6.4, and 6.0 points respectively, while average actions rose as much as 112.4% and input tokens as much as 218.3%. Simply replacing idiomatic Django names like get_*_display sent the models wandering across the repository.

Where the extra actions go

Trajectory analysis shows 81.6-83.6% of the additional actions concentrate on exploration behaviors — navigate, search, read, probe — with only about a sixth spent on editing and testing. The strategy split is telling: DeepSeek-v4-Flash already issued many actions at baseline and devoted 31.4% of its extra actions to probing, preserving relatively more of its original score; GPT-5.4-mini concentrated its extra exploration on read (41.8%) and search (33.1%) with probe at just 3.7%, leaning on static retrieval — and once familiarity was erased, its relative drop was larger.

The finding transfers to SWE-QA, a repository-level QA benchmark: GPT-5.4-mini's average score fell from 70.35 to 65.71, DeepSeek-v4-Flash slipped from 72.97 to 72.42, but its actions rose from 24.49 to 35.02 with input tokens up 59.10%.

No score drop does not mean no memorization

The most revealing result is the counterexample: on the March 2026 SWE-rebench split (110 instances created after GPT-5.4-mini's release), Pass@1 held at 17.27% under full transformation, yet interaction costs still climbed — actions +8.15%, input tokens +22.01%. For temporally held-out tasks, the authors read it as: after transformation the model needs more interaction to rebuild repository context, even when the task itself is uncontaminated.

The industry implication cuts both ways: leaderboard scores deserve a discount, since high scores may partly come from recognizing repositories; but there is no need to swing to the opposite extreme of "it's all memorization" — on uncontaminated new tasks scores did not collapse, and what rose was cost rather than a fall in capability. What should really change is evaluation itself: dynamically instantiated repository representations should become standard for coding-agent evaluation.

Paper: https://arxiv.org/abs/2609.27891
Code: https://github.com/cslsolow/Schrodinger-Repo