[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-avo-arc-agi-3-agent-harness":3,"news-related-15747718-ff24-4026-9176-433bc5553bb7":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"15747718-ff24-4026-9176-433bc5553bb7","NVIDIA AVO 刷满 ARC-AGI-3:Claude Opus 5 裸跑 30%,套上 agent 框架 100%","NVIDIA 研究团队的 agent 架构 AVO 在 ARC-AGI-3 公开集拿下 100.00 RHAE,解完 183 关;同一 Claude Opus 5 裸跑仅约 30%。核心启示:长程任务的能力是整个系统的属性,不只是模型本身。","先说一个反直觉的数字对比:同一个 Claude Opus 5,在 ARC Prize 官方跑分里大约只能拿到 30% 的成绩;而 NVIDIA 的研究团队把它装进自己的 agent 架构 AVO 之后,这个模型在 ARC-AGI-3 公开集上解完了全部 183 关,RHAE 拿到 100.00。模型一个字没改,变的只是外面那层「壳」。\n\n## 事件:AVO 刷满 ARC-AGI-3 公开集\n\n8 月 21 日,NVIDIA 在技术博客公布:其长程自主 agent 架构 AVO(Agentic Variation Operators)在 ARC-AGI-3 基准的 25 个环境、183 个关卡全部通过,RHAE 分数 100.00,总共用了 6,624 次环境动作。作为对照,此前用同一模型(Claude Opus 5)完成同样 183 关的 VISTA 系统用了 7,542 次——AVO 少用了约 12% 的动作。\n\n需要强调两点限定:这是公开集(public set)的成绩,不含半私有和私有的竞赛集;NVIDIA 自己也说明,AVO 与 VISTA 的对比不是受控消融实验,两套系统在观测表示、记忆管理、上下文管理等实现细节上都有差异。\n\n## ARC-AGI-3 测的是什么\n\nARC-AGI-3 是一个交互式推理基准:agent 进入完全陌生的游戏化环境,没有说明书、没有规则、没有目标,只能通过交互试探,推断环境动态和目标,并高效规划动作。RHAE 指标同时考核完成度和动作效率(相对首次上手的人类基线)。\n\n有意思的是 AVO 的观测方式:LLM 全程只收文本——每个观测是一个精确的 64×64 文本网格,不发送任何图像或图像 token。而 VISTA 主配置用的是渲染的 512×512 PNG。同一批任务,两种完全不同的感知通道,最后都通了。\n\n## 更早的战绩:七天自主优化 GPU 内核\n\nAVO 最初是在软件工程任务上验证的。在注意力内核优化实验里,它连续自主运行了 7 天,探索 500 多个优化方向,提交 40 个内核版本,在 NVIDIA DGX B200 系统上,演化出的多头注意力内核比 cuDNN 快最多 3.5%、比 FlashAttention-4 快最多 10.5%。之后它只花了约 30 分钟的额外自主工作,就把内核适配到了 grouped-query attention。\n\n支撑这种长程运行的是两个机制:**持久记忆**(保留历史实现、评估结果、编译器和 profiler 输出、积累的推理,让 agent 从当前状态续跑而不是反复重建搜索)和**监督者**(监测整体轨迹是否停滞、是否陷入无产出循环,必要时把主 agent 引向别的策略)。\n\n## 真正的信息:能力是系统的属性\n\nNVIDIA 的结论写得很直白:评估一个模型不等于评估一个 agent。模型能力当然重要,但决定这份能力能否转化为持续自主推进的,是外层的系统——记忆决定什么能留下来,工具决定哪些动作可行,反馈把进度锚定在现实上,恢复机制让工作能跨过单次模型调用延续下去。\n\n团队还把 AVO 接到了 GPT-5.6 Sol 上做了小规模实验:在部分匹配关卡上,Sol 壁钟时间更快,Opus 动作数更少——不同模型呈现出互补的运行画像,系统性的对比留给了后续工作。\n\n对行业的提示是:接下来一年,「agent 框架层」的工程价值可能被重新定价。大家都在卷模型参数和跑分,但把 30% 变成 100% 的,是记忆、监督、反馈循环这些看起来不起眼的系统设计。论文在 arXiv(2603.24517),原始报道见 [NVIDIA 技术博客](https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fnvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents\u002F)。\n\n所以下次看到某个模型跑分不高,先别急着下结论——它可能只是还没遇到会用它的 harness。","https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fnvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents\u002F","474eef8c-e0c3-46cf-adee-c089558220f9",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":22,"name":23,"slug":23,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"45285f33-89de-46d0-b2c9-56f83c1517b0","en","NVIDIA AVO Aces ARC-AGI-3: Claude Opus 5 Scores 30% Alone, 100% Inside the Agent Harness","NVIDIA's research agent architecture AVO scored 100.00 RHAE on the ARC-AGI-3 public set, solving all 183 levels — while the same Claude Opus 5 runs at roughly 30% on its own. The takeaway: long-horizon capability is a property of the whole system, not the model alone.","Start with a counterintuitive pair of numbers: the same Claude Opus 5 model scores roughly 30% in ARC Prize's official runs, yet after NVIDIA's research team wrapped it in their AVO agent architecture, it solved all 183 levels of the ARC-AGI-3 public set with a 100.00 RHAE score. The model didn't change at all — only the shell around it did.\n\n## What Happened: AVO Clears the ARC-AGI-3 Public Set\n\nOn August 21, NVIDIA announced on its technical blog that its long-horizon autonomous agent architecture, Agentic Variation Operators (AVO), passed all 25 environments and 183 levels of the ARC-AGI-3 benchmark with a 100.00 RHAE score, using 6,624 environment actions in total. For comparison, VISTA — an earlier system that completed the same 183 public-set levels with the same Claude Opus 5 model — used 7,542 actions, meaning AVO used roughly 12% fewer.\n\nTwo caveats matter: this covers the public set only, not the semi-private or private competition sets; and NVIDIA itself notes the AVO-vs-VISTA comparison is not a controlled ablation — the two systems differ in observation representation, memory management, context management, and other implementation details.\n\n## What ARC-AGI-3 Actually Tests\n\nARC-AGI-3 is an interactive reasoning benchmark: an agent enters completely unfamiliar game-like environments with no instructions, no stated rules, and no stated goal. It must explore through interaction, infer the environment's dynamics and objectives, and plan actions efficiently. The RHAE metric combines task completion with per-level action efficiency relative to first-time human baselines.\n\nAVO's observation channel is particularly interesting: the LLM operated entirely in text — each observation was supplied as an exact 64×64 text grid, with no images or image tokens sent to the model. VISTA's primary configuration, by contrast, used a rendered 512×512 PNG. The same set of tasks, two very different perception channels, and both eventually passed.\n\n## The Earlier Track Record: Seven Days of Autonomous GPU-Kernel Optimization\n\nAVO was first validated on software engineering work. In an attention-kernel optimization study, it ran continuously and autonomously for seven days, explored more than 500 optimization directions, and committed 40 kernel versions. On NVIDIA DGX B200 systems, the evolved multihead attention kernels outperformed cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5%. The agent then adapted the evolved kernel to grouped-query attention in roughly 30 minutes of additional autonomous work.\n\nTwo mechanisms sustain this kind of long-horizon run: **persistent memory** (carrying forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, so the agent resumes from its current state instead of rebuilding the search) and a **supervisor** (monitoring the broader trajectory for stagnation or unproductive loops and redirecting the main agent when needed).\n\n## The Real Message: Capability Is a Property of the System\n\nNVIDIA's conclusion is blunt: evaluating a model is not the same as evaluating an agent. Model capability matters enormously, but the surrounding system determines whether that capability converts into sustained autonomous progress — memory decides what survives, tools determine which actions are possible, feedback grounds progress in reality, and recovery mechanisms let work continue beyond a single model invocation.\n\nThe team also paired AVO with GPT-5.6 Sol on a challenging subset of games: in those limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions — complementary operating profiles across models, with a systematic comparison left to future work.\n\nThe industry implication: over the next year, the engineering value of the agent-framework layer may get repriced. Everyone is competing on model parameters and raw benchmark scores, but what turned 30% into 100% was memory, supervision, and feedback loops — system design that looks unglamorous. The paper is on arXiv (2603.24517); the original report is on the [NVIDIA Technical Blog](https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fnvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents\u002F).\n\nSo the next time a model posts an unimpressive score, hold your judgment — it may simply not have met a harness that knows how to use it.","nvidia-avo-arc-agi-3-agent-harness","2026-08-23T15:30:00Z","2026-08-23T15:11:08.711636Z","2026-08-23T15:11:08.711648Z",true,"agent",74,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"32b938b6-01a3-43c9-b040-14db6c5f57c6","NVIDIA 把 Agent 装进一个 Python 类:被忽略的 NOOA,一半 token 跑出 SWE-bench 82.2%","nvidia-nooa-python-agent-framework","2026-08-23T17:20:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"0237222a-602b-47ef-9431-468009904428","FACET 先建环境再写任务:1.2K 轨迹把 Qwen3.5-27B 推到 Terminal-Bench 47.57,逼近 397B","facet-terminal-task-synthesis","2026-08-19T06:19:20+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"454f9530-20d7-428c-82d9-9175fa5b883a","Claude 推黎曼 zeta 下界到 67.2%：60 subagent + Lean","claude-zeta-bound-67-percent-multi-agent-lean","2026-08-17T07:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"7d371b09-9792-465d-b73a-3d0af4735129","InferenceBench：15 个前沿 Agent 自主做 LLM 推理优化","inferencebench-open-ended-llm-optimization","2026-08-16T12:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"3c6fcf46-f5bb-4136-931c-69cd64216e12","Skill-Use 基准揭示 Agent 短板：会做任务，不等于会用 Skill","skill-use-agent-harness-benchmark","2026-08-06T08:00:00+00:00"]