[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-facet-terminal-task-synthesis":3,"news-related-0237222a-602b-47ef-9431-468009904428":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0237222a-602b-47ef-9431-468009904428","FACET 先建环境再写任务:1.2K 轨迹把 Qwen3.5-27B 推到 Terminal-Bench 47.57,逼近 397B","中科大等机构提出 FACET 终端任务合成框架:先建环境再生成任务,1.2K 条轨迹微调让 Qwen3.5-27B 在 Terminal-Bench 2.1 达 47.57,逼近 397B 基座,任务与模型已开源。","终端 Agent 的训练一直有个绕不开的瓶颈：要让模型学会操作真实系统，你需要大量\"可执行\"的训练任务——每个任务至少耦合四件东西：一条指令、一个初始化环境、一份参考解法、一个可执行验证器。这四样如果是从彼此不一致的假设里分别生成的，产出的任务要么根本无解，要么验证器判错。中科大、上海人工智能实验室与复旦的团队在 8 月 19 日提交到 arXiv 的 FACET 框架（[arXiv:2608.18580](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.18580)），给出的答案朴素得近乎反直觉：先把世界建好，再写任务。\n\n## 任务合成的信息流失问题\n\n论文指出的第一个痛点是信息保真。多阶段合成管线在逐步转换的过程中，会把原始素材里编码的目标、依赖关系、状态迁移和过程性约束一点点丢掉——最后生成的任务表面上完整，实际上和源头的意图已经对不上。第二个痛点是跨组件一致性：指令、解法、验证器各自生成，谁也不知道对方假设的环境长什么样。\n\nFACET（Fine-grained Agentic Construction of Executable Tasks）的解法分三步走：先从信息源中获取并组织出 71K 以上可复用的 agent 技能，构建\"场景—技能\"库；再重建场景，恢复目标、依赖、中间状态、工具和输入输出契约；最后才是关键的一步——先把执行环境真正跑起来并修复，让容器状态成为指令、解法、验证器三者共享的接地（grounding）。哪个组件验证失败就定点修复哪个，不推倒重来。项目页把这套哲学压缩成一句话：Build the world first, then write the task。\n\n## 1.2K 条轨迹撬动的提升\n\n结果数字相当扎实。在 Terminal-Bench 2.1 上，仅用 FACET 任务收集到的 1.2K 条成功轨迹做微调：Qwen3.5-4B 从 17.60 提到 24.72（+7.12），Qwen3.5-9B 从 27.34 提到 35.58（+8.24），Qwen3.5-27B 从 40.82 提到 47.57（+6.75）。最有意思的参照是：27B 微调后的 47.57，距离 397B-A17B 基座在同一评测设定下的 49.06 只差 1.49 分——模型尺寸约是后者的十五分之一。需要说明，这些是论文自测数据，尚无第三方复现。\n\n更值得看的是消融。任务生成的顺序本身决定了任务有效率：环境→指令→解法→验证器的正向顺序，最终验证通过率达到 83%；联合生成只有 65%，反向生成 63%。首轮通过率差距更悬殊（46.5% 对 37.5% 和 24.2%）。这组对照直接回答了\"到底是数据多还是顺序对\"的问题——在同样的模型和算力下，把环境放在生成链的最前面，本身就是任务有效性的来源。构造漏斗也印证了定向修复的价值：7,852 个场景种子，7,504 个环境构建成功（95.7%），首轮通过验证的任务只有 2,856 个（38.35%），经组件级修复后最终达到 6,078 个（81.63%）——修复环节让产出翻了一倍还多。\n\n## 数据效率视角\n\n把 FACET 放进同脚手架（Terminus-2）的数据集对比里看，轨迹效率的优势很直观：其他终端 Agent 数据集普遍需要 5K 到 32K 条轨迹（Nemotron-Terminal 5K、Terminal-Lego 32K），FACET 用 1.2K 条就支撑起 6,078 个任务、平均每任务 22.77 个可执行测试的密度。论文附录还给出轨迹行为统计：95.6% 的轨迹首回合是纯观察、84.7% 的回合包含观察命令——这种\"先看再做\"的行为模式正是从环境接地里自然长出来的。\n\n## 所以呢\n\nFACET 开源了 6,020 个公开发布的任务（FACET-Terminal-Tasks-6k）和 4B、9B、27B 三个微调 checkpoint，登上 Hugging Face Daily Papers 榜单（111 票）。对做 Agent 训练的团队，这篇工作的可迁移结论不是某个具体数字，而是一条工程原则：合成数据的下一步不是造得更多，而是造得更可执行——环境先落地，任务才站得住。当所有人都在卷模型和算力时，把\"任务本身是否真实可解\"当作一等公民来工程化，可能才是数据侧最被低估的杠杆。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.18580","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"6c98bb2b-462d-4e98-a4ab-6344665621af","en","FACET: Build the Environment First, Then Write the Task","USTC and Shanghai AI Lab propose FACET: realize the environment before generating tasks. 1.2K trajectories lift Qwen3.5-27B to 47.57 on Terminal-Bench 2.1.","Training terminal agents has a persistent bottleneck: to teach models to operate real systems, you need large volumes of executable training tasks — and each task couples at least four artifacts: an instruction, an initialized environment, a reference solution, and an executable verifier. If these artifacts are generated from inconsistent assumptions, the resulting task is either unsolvable or incorrectly evaluated. A team from USTC, Shanghai AI Laboratory, and Fudan University posted their FACET framework to arXiv on August 19 ([arXiv:2608.18580](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.18580)), with an answer that sounds almost counter-intuitively simple: build the world first, then write the task.\n\n## The information-loss problem in task synthesis\n\nThe paper identifies two pain points. First, information preservation: multi-stage synthesis pipelines gradually discard the goals, dependencies, state transitions, and procedural constraints encoded in original sources — the final task looks complete but no longer matches the source intent. Second, cross-artifact consistency: instruction, solution, and verifier are generated separately, none aware of what environment the others assume.\n\nFACET (Fine-grained Agentic Construction of Executable Tasks) proceeds in three steps: acquire and organize 71K+ reusable agent skills into a scenario-skill repository; reconstruct scenarios by recovering goals, dependencies, intermediate states, tools, and I\u002FO contracts; and finally the key move — realize and repair the execution environment first, making the container state the shared grounding for instruction, solution, and verifier. When a component fails validation, only that component is repaired, not the whole pipeline. The project page compresses this philosophy into one line: build the world first, then write the task.\n\n## What 1.2K trajectories buy\n\nThe numbers are solid. On Terminal-Bench 2.1, fine-tuning with just 1.2K successful trajectories collected from FACET tasks lifts Qwen3.5-4B from 17.60 to 24.72 (+7.12), Qwen3.5-9B from 27.34 to 35.58 (+8.24), and Qwen3.5-27B from 40.82 to 47.57 (+6.75). The most interesting reference point: the fine-tuned 27B reaches 47.57, within 1.49 points of the Qwen3.5-397B-A17B base (49.06) under the same evaluation setting — at roughly 1\u002F15 the model size. Note these are the authors' own measurements; no independent replication exists yet.\n\nThe ablations deserve more attention than the headline. Generation order alone determines task validity: the forward order — environment, instruction, solution, then solution-aware verifier — reaches an 83% final yield, versus 65% for joint generation and 63% for reverse. First-pass validity gaps are wider still: 46.5% versus 37.5% and 24.2%. This comparison directly answers the question of whether it is data volume or ordering that matters — under the same models and compute, putting the environment first in the generation chain is itself a source of task validity. The construction funnel confirms the value of targeted repair: 7,852 scenario seeds yielded 7,504 successful environments (95.7%), but only 2,856 first-pass valid tasks (38.35%); component-level repair brings the final count to 6,078 (81.63%) — more than doubling the output.\n\n## The data-efficiency lens\n\nPlaced in the same-scaffold (Terminus-2) dataset comparison, FACET's trajectory efficiency is striking: other terminal-agent datasets typically require 5K to 32K trajectories (Nemotron-Terminal 5K, Terminal-Lego 32K), while FACET's 1.2K trajectories support 6,078 tasks with an average of 22.77 executable tests per task. The appendix also reports trajectory behavior statistics: 95.6% of trajectories open with an observation-only first turn, and 84.7% of turns include observation commands — a look-before-acting pattern that grows naturally out of environment grounding.\n\n## So what\n\nFACET open-sources 6,020 public-release tasks (FACET-Terminal-Tasks-6k) and three fine-tuned checkpoints at 4B, 9B, and 27B, and has climbed the Hugging Face Daily Papers board (111 upvotes). For teams training agents, the transferable lesson here is not any single number but an engineering principle: the next step for synthetic data is not producing more, but producing what actually executes — land the environment first, and the task stands on solid ground. While everyone competes on models and compute, treating \"is this task actually solvable\" as a first-class engineering concern may be the most underrated lever on the data side.","facet-terminal-task-synthesis","2026-08-19T06:19:20Z","2026-08-22T15:09:08.330268Z","2026-08-22T15:09:08.330283Z",true,"agent",47,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"32b938b6-01a3-43c9-b040-14db6c5f57c6","NVIDIA 把 Agent 装进一个 Python 类:被忽略的 NOOA,一半 token 跑出 SWE-bench 82.2%","nvidia-nooa-python-agent-framework","2026-08-23T17:20:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"c4acfb32-44d4-45a9-ac27-ec9fff4e1eca","Mistral 开源 Leanstral 1.5:6B 激活参数刷新形式化推理 SOTA","mistral-leanstral-1-5","2026-07-04T00:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"90af7ff5-b985-42d5-97c7-63a9579b7527","VitaBench 2.0：给 LLM Agent 出「长期用户建模」考卷，SOTA 也不及格","vitabench-2-0-long-term-user-modeling","2026-06-25T14:01:00+00:00"]