[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-apodex-1-1-agentic-execution-pivot-rl":3,"news-related-f26ace13-9c96-47ea-a528-b6682a22aa1e":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","Apodex 8 月 24 日发布 1.1 版:把推理从生成报告搬进真实任务执行,异步 Agent Team 加中途介入撑起长任务;训练侧 PIVOT-RL 在数十万条轨迹中定位关键决策点做局部学习;35B mini 与开源 FrontierAgent harness 支持本地部署。","过去两年,\"深度研究\"类产品的通用形态是:AI 搜网页,整理成结构化报告。但真实科研工作往往在报告写完才真正开始——原始输入是几十篇论文、实验记录和各种专业格式文件,要清洗数据、写代码、跑分析,计划随中间结果不断调整。Apodex 8 月 24 日发布的 1.1 版模型正冲着这个缺口:把推理从报告里拿出来,放进真实任务的执行过程。\n\n## PIVOT-RL:在数十万条轨迹里找关键决策点\n\n长轨迹任务的强化学习有个核心难题:终局结果只提供粗粒度监督——成功的轨迹里仍藏着低效的中间决策,失败的轨迹前半段也有真正有用的工作。Apodex 的解法叫 PIVOT-RL:对数十万条轨迹做事后分析,定位改变任务走向的关键决策点,保留有效前缀,构造带短纠正提示的局部继续任务;提示只在训练时提供方向引导,推理时不存在。整体采用 SFT 加 agentic RL 组合,让失败恢复和任务交付成为模型学到的行为。能力沿两条路径扩展:Environment Scaling(文件、搜索、代码等环境的数千万训练任务)与 Agentic Coordination Scaling(任务分解与多 Agent 协调),共用叫 AgentOS 的运行时基座。\n\n## 异步 Agent Team 与 Statement Review\n\nDeep Discover 模式下,模型动态组建异步 Agent Team——不是预写编排脚本,而是模型自己判断任务怎么分解、交给几个 Subagent,多个 Subagent 并行探索,中间结果持续回喂主任务。用户中途介入被当作任务的一部分:系统判断哪些中间结果仍有效,从当前状态继续而非推倒重来。交付前的 Statement Review 把生成与审查分开:某条数据是否真支持统计显著性、被引论文是否真说了报告转述的话,关键结论先过独立检查才到用户手里。\n\n## 三个带数字的用例\n\n法律场景的破产偏颇清偿案中,系统重建六笔付款的历史,得出 55 万美元净敞口,与参考答案一致,并给出 17.5 万到 35 万美元的和解区间与法律备忘录。金融场景的外汇对冲任务里,系统算出 760 万美元领式组合净成本和 1.1172 的盈亏平衡汇率,还发现 CFO 对会计处理的假设有误,按 ASC 815 重新推导了影响。科研场景最硬核:从 7M6J 蛋白质结构出发,用 Martini 3 粗粒化方法生成力场拓扑,搭出包含 69,652 个模拟位点的 GROMACS 模拟体系并完成能量最小化。\n\n## 开源部分与 AI4AI\n\n本地部署方面,35B 的 Apodex 1.1 mini 公开参数规模;配套 FrontierAgent 执行框架已在 GitHub 开源,命令行 TUI 支持 ReAct 与 Agent Team 两种模式,macOS 和 Linux 一条命令开箱即跑,不依赖 Docker。AI4AI 实验则让 Apodex 1.1 全程当老师——自动出题、过滤轨迹、评估迭代,无人工参与也不依赖其他强模型——用它训练 Qwen3.5-0.8B 在 200 道检索注释题上,10 轮迭代后总分从 51.0% 升到 56.0%。\n\n## 所以呢\n\n这家公司画的终局叫 Heavy-Duty Solver:接下越来越复杂的长周期工作,交付端到端可验证的结果,下一代 2.0 将从预训练阶段从头构建。当\"会答题\"和\"能把一件事干完且经得起核查\"分成两种能力,评测重心也会从单轮问答搬到长轨迹执行——Apodex 自建的两个基准方向就在这里(官方发布原文:https:\u002F\u002Fwww.apodex.com\u002Fblog\u002Fapodex-1.1-scaling-agentic-intelligence-for-complex-work)。","https:\u002F\u002Fwww.apodex.com\u002Fblog\u002Fapodex-1.1-scaling-agentic-intelligence-for-complex-work","acf51543-e748-4b67-a08c-1885e3100b5b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"012658cd-0709-40f0-8b7a-bf630a0c9cad","en","Apodex 1.1 Moves Reasoning Into Execution: PIVOT-RL, Open 35B Mini","Released August 24, Apodex 1.1 shifts reasoning from report generation into real task execution, with an asynchronous Agent Team, mid-task user intervention, and independent Statement Review. Training-side PIVOT-RL localizes consequential decision points across hundreds of thousands of trajectories, while the 35B mini and the open-source FrontierAgent harness support local deployment.","For the past two years, the standard shape of \"deep research\" products has been an AI that searches a large pile of web pages and assembles a structured report. But real research work usually only begins after the report is written. The raw input may be dozens of papers, lab records, tables, images, and files in specialized formats; researchers must clean data, choose methods, write code, run analyses, and interpret results. A single task contains multiple interdependent stages, and the plan keeps adjusting based on intermediate results. Apodex 1.1, released on August 24, is built for exactly this gap. In the company's own one-liner: it takes reasoning out of the report and puts it into the execution of real tasks.\n\n## Two Scaling Paths and One Runtime Foundation\n\nCapability grows along two complementary paths. Environment Scaling expands the executable environments the model can learn and act in — files, search, code, and other environment categories — with training tasks at a scale of tens of millions. Agentic Coordination Scaling expands a task's ability to be decomposed, coordinated, integrated, and reorganized. Both paths run on a shared runtime foundation called AgentOS, which maintains tool calls, file state, and task progress during a single task, manages the creation, scheduling, and lifecycle of Subagents, and provides a unified verification mechanism. Training combines SFT with agentic RL: SFT gives the model the basic shape of behaviors like tool calling, task decomposition, and multi-agent coordination, and agentic RL then refines these on real execution and coordination trajectories, so failure recovery and task delivery become learned behaviors rather than effects produced by external orchestration.\n\n## PIVOT-RL: Finding Consequential Decisions in Hundreds of Thousands of Trajectories\n\nReinforcement learning on long-horizon tasks has a core problem: terminal outcomes provide only coarse supervision. A successful trajectory can still contain inefficient or weakly grounded intermediate decisions, and a failed one can still contain genuinely useful early work. Apodex's answer is PIVOT-RL. Using Hindsight-Guided Trajectory Localization, it runs retrospective analysis over a training corpus of hundreds of thousands of trajectories and questions to identify the consequential decision points — the pivots — where the model starts following an unproductive strategy, relies on insufficient evidence, misuses a tool, or fails to revise a wrong assumption. At each pivot, the working prefix is preserved and a localized continuation task is constructed with a short corrective hint. The hint only provides directional guidance during training, is never a prediction target, and is absent at inference time. For stateful tasks, the corresponding executable environment state is restored. Localized continuations are mixed with full, unhinted tasks during training, so the model learns efficiently where it is genuinely prone to error while retaining the ability to solve a complete task on its own.\n\n## Asynchronous Agent Team and Statement Review\n\nIn Deep Discover mode, the model dynamically organizes an asynchronous Agent Team based on the task — not a pre-written orchestration script, but the model itself deciding whether a task can be decomposed, how, and across how many Subagents, continually deciding when to consolidate results. Multiple Subagents explore different subtasks or candidate hypotheses in parallel, feeding intermediate results back to the main task continuously instead of waiting for every branch to finish. Mid-task user intervention is treated as part of the task itself: the system must understand the new requirement, judge which completed intermediate results remain valid, update the plan, and continue from the current state rather than starting over. Before delivery, Statement Review keeps generation and review as distinct steps: key claims — whether a piece of data genuinely supports statistical significance, whether a cited paper actually says what the report claims, whether a computed result matches what the code produced — go through an independent check before reaching the user.\n\n## Three Use Cases With Numbers\n\nIn a legal preference-liability case from a corporate bankruptcy, the system reconstructed the timing and collection history of six payments, arrived at a net exposure of $550K matching the reference answer, and further provided a settlement range of $175K–$350K plus a client-ready legal memo. In a cross-border FX hedging task, it worked out option premiums under different structures, payoffs across five settlement exchange rates, a $7.6M net collar cost, and a 1.1172 breakeven rate — and caught that the CFO's underlying assumption about accounting treatment was itself wrong, re-deriving the impact of all three structures under ASC 815 on P&L, OCI, and reclassification risk. The research case is the most hardcore: starting from the 7M6J protein structure, it chose the Martini 3 coarse-grained method to generate force-field topologies, replicated three copies of the protein, added water and physiological-concentration salt ions, and built a complete GROMACS simulation system with 69,652 simulation sites, completing energy minimization.\n\n## The Open-Source Part and AI4AI\n\nOn the local ecosystem side, the 35B Apodex 1.1 mini offers a locally deployable option with a disclosed parameter count; the company says it reaches the performance band of selected frontier systems on professional work, finance, and scientific research. The companion FrontierAgent execution harness is open-sourced on GitHub, with a native command-line TUI supporting both ReAct single-agent and Agent Team modes, running out of the box with a single command on macOS and Linux without depending on Docker. The AI4AI experiment shows another use: Apodex 1.1 acts as the Teacher — generating questions, filtering correct trajectories, evaluating, and iterating with no human involvement and no reliance on other strong models — and was used to train Qwen3.5-0.8B on three task types covering 200 questions in clinical trial and drug information retrieval, protein structure database lookup, and protein sequence and function annotation. After 10 rounds of automated iteration, the small model's overall score rose from 51.0% to 56.0%.\n\n## So What\n\nThe end state this company draws for itself is the Heavy-Duty Solver: a system that takes on increasingly complex, long-horizon work while delivering results that remain verifiable end to end, with Apodex 2.0 to be built from the pretraining stage up. The signal for practitioners is direct: as \"answering questions\" and \"seeing a piece of work through to completion, verifiably\" split into two different capabilities, evaluation will shift from single-turn Q&A to long-trajectory execution — Apodex has built two benchmarks of its own, FrontierSearchBench and FrontierResearchBench, pointing exactly in that direction (official release: https:\u002F\u002Fwww.apodex.com\u002Fblog\u002Fapodex-1.1-scaling-agentic-intelligence-for-complex-work).","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00Z","2026-08-25T15:23:21.570156Z","2026-08-25T15:23:21.570166Z",true,"agent",55,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"79c1684f-f61d-4799-b3d0-6450c4ad10e8","Muse Glimmer:Meta 把 30B 「常驻本地的智能体」开源,把 Agent 拉到笔记本里 7×24 跑","muse-glimmer-30b-open-agentic-local","2026-08-10T00:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"5ef8c5da-7221-4e22-bd46-c1ef55c3120d","韩国 Solar Open 2：Linear Attention 进 MoE，250B\u002F15B 激活","upstage-solar-open-2","2026-07-23T03:30:00+00:00"]