[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-schrodinger-repo-swe-bench-memorization":3,"topics-all":38,"news-related-448dbda7-e4ac-44ba-8a40-7af01012b8df":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"448dbda7-e4ac-44ba-8a40-7af01012b8df","把仓库\"换脸\"再测:SWE-bench 高分最多掉 14.4 分,编码 agent 疑在背题","上海交大团队提出 SchrodingerRepo 框架:评测时把 SWE-bench 的测试仓库动态换脸——重写问题描述、命名空间改名、重排布局、重写函数体,行为不变但熟悉感全失。四个模型 Pass@1 集体下滑 6.0-14.4 个百分点,新增动作八成花在重新认路。代码已开源。","SWE-bench 榜单上的高分,究竟多少来自真实推理,多少来自\"认得这个仓库\"?上海交大团队 8 月 21 日提交到 arXiv 的论文(arXiv:2609.27891)给出了一种把两者分开的实验设计:不换题目,只把测试仓库\"换脸\",看 agent 还能考多少分。\n\n## 评测时才\"坍缩\"的仓库\n\n论文提出 SchrodingerRepo 框架,核心思想借自薛定谔的猫:测试仓库被当作评测时的潜变量,只在 agent 进入评测环境那一刻才动态实例化。实例化保留原始可执行行为,但通过四个层级逐步抹掉熟悉线索——重构问题描述、命名空间改名、文件内布局重排、功能保持的代码重写。作者来自上海交大和西安交大,代码以 MIT 协议开源。\n\n动机实验先给出泄漏证据:对每个被测模型,超过 65% 的 SWE-bench Verified 题目有明确数据泄漏,超过 18% 的题目能被回忆到补丁或测试级别。这呼应了 OpenAI 此前\"SWE-bench Verified 已无法可靠衡量前沿编码能力\"的判断。\n\n## 四个模型集体掉分\n\n团队用 mini-swe-agent 脚手架在 SWE-bench Verified 上评测四个后端:GPT-5.4-mini、GPT-5.1、Gemini-3.1-Flash-Lite、DeepSeek-v4-Flash。四层变换全开后,Pass@1 全线下滑:GPT-5.4-mini 从 46.8% 降到 35.6%,DeepSeek-v4-Flash 从 72.8% 降到 66.8%,GPT-5.1 从 44.6% 降到 36.2%,Gemini-3.1-Flash-Lite 从 56.7% 降到 42.3%,整体降 6.0-14.4 个百分点(p\u003C0.05)。\n\n最狠的不是重写代码,而是命名空间改名:单这一层就让三个模型分别掉 7.4、6.4、6.0 个百分点,平均动作数最高涨 112.4%,输入 token 最高涨 218.3%。仅仅把 Django 的 `get_*_display` 这类惯用命名换掉,模型就开始满仓库乱撞。\n\n## 多花的动作都花在哪了\n\n轨迹分析显示,新增动作的 81.6-83.6% 集中在 navigate、search、read、probe 这类探索行为,只有一成多用于编辑和测试。策略也在分化:DeepSeek-v4-Flash 新增动作中 31.4% 是 probe(轻量运行时试探),靠多试探保住了更多原始分数;GPT-5.4-mini 更依赖 read 和 search 静态检索,熟悉感一被抹掉,相对降幅反而更大。\n\n结论可迁移到仓库级问答基准 SWE-QA:GPT-5.4-mini 平均分从 70.35 跌到 65.71,DeepSeek-v4-Flash 仅微降,但后者的动作数从 24.49 涨到 35.02,输入 token 增 59.10%。\n\n## 分数没掉,不代表没背题\n\n最能说明问题的是反例实验:在 2026 年 3 月 SWE-rebench 分片(110 个 GPT-5.4-mini 发布后新造的题)上,Pass@1 在全量变换下保持 17.27% 不变,但交互成本照涨——动作 +8.15%,输入 token +22.01%。作者解读:变换后模型需要更多交互来重建仓库上下文,即便题目本身没有泄漏。\n\n对行业的含义是双向的:榜单分数要打折看,高分可能部分来自\"认得仓库\";但也不必走向\"全是背题\"的极端——没泄漏的新题上分数并未崩塌,涨的是成本而非降的能力。真正该变的是评测本身:动态实例化的仓库表示,应该成为编码 agent 评测的标配。\n\n论文:arxiv.org\u002Fabs\u002F2609.27891\n代码:github.com\u002Fcslsolow\u002FSchrodinger-Repo","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.27891","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"e45bc728-8681-4a1f-ab2b-40935e585f0c","en","Swap the repo's face and SWE-bench scores drop up to 14.4 points","SJTU's SchrodingerRepo rewrites SWE-bench test repos at eval time; Pass@1 drops 6.0-14.4pp across four models, 80% of extra actions spent re-exploring.","How much of a high SWE-bench score comes from real reasoning, and how much from simply recognizing the repository? A paper submitted to arXiv on Aug 21 (arXiv:2609.27891) by a Shanghai Jiao Tong University-led team offers an experiment that separates the two: keep the tasks, but swap the face of the test repository, and see how much the agent still scores.\n\n## A repository that collapses only at evaluation time\n\nThe paper proposes SchrodingerRepo, whose core idea borrows from Schrödinger's cat: the test repository is treated as an evaluation-time latent variable, dynamically instantiated only when the agent enters the evaluation environment. The instantiated version preserves the original executable behavior while progressively eroding familiar cues through four transformation levels — problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting. Authors come from SJTU (co-first authors Silin Chen and Yufei Yang, corresponding author Xiaodong Gu) and Xi'an Jiaotong University; the code is open-sourced under MIT.\n\nA motivation experiment first quantifies the leakage: for every evaluated model, more than 65% of SWE-bench Verified instances show clear data-leakage evidence, and more than 18% can be recalled at the patch or test level. This echoes OpenAI's earlier judgment that SWE-bench Verified no longer reliably measures frontier coding capability.\n\n## Four models drop together\n\nUsing the mini-swe-agent scaffold on SWE-bench Verified, the team evaluated four backends: GPT-5.4-mini, GPT-5.1, Gemini-3.1-Flash-Lite, and DeepSeek-v4-Flash. With all four transformation levels enabled, Pass@1 fell across the board: GPT-5.4-mini from 46.8% to 35.6%, DeepSeek-v4-Flash from 72.8% to 66.8%, GPT-5.1 from 44.6% to 36.2%, and Gemini-3.1-Flash-Lite from 56.7% to 42.3% — an overall drop of 6.0-14.4 percentage points, statistically significant (p\u003C0.05).\n\nThe harshest level was not code rewriting but namespace remapping: that single level cut three models by 7.4, 6.4, and 6.0 points respectively, while average actions rose as much as 112.4% and input tokens as much as 218.3%. Simply replacing idiomatic Django names like `get_*_display` sent the models wandering across the repository.\n\n## Where the extra actions go\n\nTrajectory analysis shows 81.6-83.6% of the additional actions concentrate on exploration behaviors — navigate, search, read, probe — with only about a sixth spent on editing and testing. The strategy split is telling: DeepSeek-v4-Flash already issued many actions at baseline and devoted 31.4% of its extra actions to probing, preserving relatively more of its original score; GPT-5.4-mini concentrated its extra exploration on read (41.8%) and search (33.1%) with probe at just 3.7%, leaning on static retrieval — and once familiarity was erased, its relative drop was larger.\n\nThe finding transfers to SWE-QA, a repository-level QA benchmark: GPT-5.4-mini's average score fell from 70.35 to 65.71, DeepSeek-v4-Flash slipped from 72.97 to 72.42, but its actions rose from 24.49 to 35.02 with input tokens up 59.10%.\n\n## No score drop does not mean no memorization\n\nThe most revealing result is the counterexample: on the March 2026 SWE-rebench split (110 instances created after GPT-5.4-mini's release), Pass@1 held at 17.27% under full transformation, yet interaction costs still climbed — actions +8.15%, input tokens +22.01%. For temporally held-out tasks, the authors read it as: after transformation the model needs more interaction to rebuild repository context, even when the task itself is uncontaminated.\n\nThe industry implication cuts both ways: leaderboard scores deserve a discount, since high scores may partly come from recognizing repositories; but there is no need to swing to the opposite extreme of \"it's all memorization\" — on uncontaminated new tasks scores did not collapse, and what rose was cost rather than a fall in capability. What should really change is evaluation itself: dynamically instantiated repository representations should become standard for coding-agent evaluation.\n\nPaper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.27891\nCode: https:\u002F\u002Fgithub.com\u002Fcslsolow\u002FSchrodinger-Repo","schrodinger-repo-swe-bench-memorization","2026-09-24T17:11:54Z","2026-09-24T17:12:06.817435Z","2026-09-24T17:12:06.817443Z",true,"agent",524,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c4acfb32-44d4-45a9-ac27-ec9fff4e1eca","Mistral 开源 Leanstral 1.5:6B 激活参数刷新形式化推理 SOTA","mistral-leanstral-1-5","2026-07-04T00:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"e4b3903e-c65f-46cb-90ae-81502eb8cdd9","CliffCompaction开源:只删不改的会话压缩,长程Agent成本砍半","cliffcompaction-truncate-only-agent-compaction","2026-09-23T17:10:56+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00"]