[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-upstage-solar-pro-4-agent-reliability-closed-llm":3,"news-related-e75069c6-f15c-4ff9-8b11-404d705442e8":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e75069c6-f15c-4ff9-8b11-404d705442e8","Upstage Solar Pro 4:把「agent 跑得稳」做成新一代闭源模型卖点","Upstage 8 月 11 日上线 Solar Pro 4 闭源商用模型:512K 上下文,综合分 42,Terminal-Bench v2.1 57、AA-LCR 71。定价 0.30\u002F1.20 美元每百万 token,主打企业级 agent 可靠性。","韩国 AI 公司 Upstage 在 8 月 11 日上线了 Solar Pro 4 闭源商用大模型。这一代不再主打「更大、更强」的通用能力分数,而是把「agent 跑得稳」作为主轴——围绕长文档、终端任务、多轮工具调用,以及拒绝编造这类企业级问题重做了训练和评估。\n\n## 一组 agent 任务分数\n\nSolar Pro 4 上下文窗口 512K,输出上限 128K,支持英语、韩语、日语三种语言,默认开启推理。Upstage 在官方博客中给出的关键 agent 任务分数,基于 Artificial Analysis 的基准体系:\n\n- Terminal-Bench v2.1(真实 shell 里完成多步任务):57\n- τ³-Banking(多轮工具调用找到正确策略):23\n- AA-LCR(跨约 10 万 token 长文档做综合推理):71\n\nUpstage 自己对照 Solar Open 2 给的同口径对比是:Terminal-Bench v2.1 +13.8、BrowseComp +11.9、GDPval-AA v2 +7.4、AA-LCR +8.3;在 GPQA Diamond、MMLU-Pro、LiveCodeBench、AIME 2026 等通用知识与数学代码基准上,两个模型差距很小。Artificial Analysis 给出的综合指数是 42 分——Upstage 自己的话是「比 Solar Pro 3 高出三倍多,超过 Nvidia Nemotron 3 Ultra 的 38、Google Gemini 3.5 Flash-Lite 的 37」。\n\nThe New Stack 在 8 月 20 日的独立报道里,转述了 Upstage 美国 CEO Kasey Roh 的一句话:「把前沿模型留给前沿问题;我们做的是拉车的马」。她给的成本对照是:300K 输入 + 15K 输出的文档事实核查任务,按列表价算,「前沿模型每任务约 1 美元,Solar Pro 4 约 0.10 美元」。\n\n## 训练侧的核心:用 OfficeVerse 合成「已完成的工作」\n\nSolar Pro 4 的训练数据由 Upstage 自研的 OfficeVerse 流水线生成——从真实公开数据合成 11 个行业域、12 类任务型办公工作,以最终交付物是否「pass or fail」作为唯一打分标准。Upstage 自己说,Solar Pro 4 的训练和验证用「和真实工作形状一致」的任务。同一套流水线也产出了 Ko-GDPval——一个韩国办公工作基准。\n\n这套思路的对照面是「刷榜式训练」——很多闭源模型在公开基准上冲到很高,但面对真实办公流时会出现「表格拉到第 200 行之后开始糊弄,把单元格用前面行的值模式匹配填上」这种 silent failure。Roh 在 The New Stack 的采访里点名的就是这个失败模式:模型先正常读到 200 行,之后开始略过行、用前面行的模式填充,输出格式还是好的,直到最后对账的人手动核对才被发现。「这是我们专门训练去避免的失败模式」。\n\n## 「找不到依据就说不」是这一代的核心交付物\n\nSolar Pro 4 在文档问答中显式区分三种判定:grounded(条款存在并给引用)、not-in-document(条款不存在)、mismatch(两份文档中数字不一致,标记为差异)。模型在「无法验证」时直接说「无法验证」,而不是用语言模式生成一个看起来对的答案——博客里那张示例图里,模型只回答了问题里一半能找到的内容,另一半明确告知「文档中没有提供」。\n\n这件事对企业的实际意义大于「又多了一个 SOTA」:在金融、法律、保险、制造业、供应链这些受监管行业里,模型生成一个看起来漂亮但找不到来源的答案,下游的所有工作都会被污染。Solar Pro 4 走的是「找到的依据给引用,找不到就标记」,把可追溯性变成默认行为。\n\n## 价格与可用性\n\nSolar Pro 4 在 Upstage Console、OpenRouter、Hermes Agent(Nous Research 的 agent 环境)、Upstage Studio 上同时上线,定价是:输入 0.30 美元\u002F百万 token,缓存输入 0.06 美元\u002F百万 token,输出 1.20 美元\u002F百万 token。Upstage 在 8 月 11 日到 9 月 10 日 23:59 UTC 之间给出 9 折列表价的发布促销。\n\nThe New Stack 引述 Upstage 提供的数据:Solar Pro 4 在 OpenRouter 上线一周内,token 消耗超过 370B。Upstage 的既有合作伙伴是 AWS 和 AMD;公司 2025 年起在 San Jose 设立美国总部。\n\n## 一段判断\n\nSolar Pro 4 不是一个「比 Sonnet\u002FGPT 更大」的模型,也不是为了把基准表刷得更亮。它押注的是企业里真正部署 agent 时遇到的那类失败——长文档中段悄悄糊弄、表格数据用前文模式填空、引用生成不存在的条款。把这三类问题从「agent 默认行为」里拆出来,配以更便宜的单价,Upstage 在做的事本质上是「agent 部署的工程化」,而不是「模型能力军备竞赛」的下一轮。对国内外的开发者来说,这都是值得停下来看一眼的方向:在 SWE-Bench、Terminal-Bench、AA-LCR 这些 agent 任务基准刷到顶之后,真正决定能不能在企业里跑起来的,是模型在几百行表格、几十页合同、上百次工具调用里保持「silent failure 接近 0」的能力。","https:\u002F\u002Fwww.upstage.ai\u002Fblog\u002Fen\u002Fsolar-pro-4","895e85f1-ff1b-444d-979b-62f10a6be5a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"62a0408c-027e-4515-807d-5f19dc5e1390","korean",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"7ef4f682-ad80-404d-bd1f-5f1397b7fe8e","en","Upstage Solar Pro 4: agent reliability as the new closed-LLM pitch","Upstage shipped Solar Pro 4 on Aug 11, 2026: a 512K-context closed LLM scored 42 on the Artificial Analysis index, with Terminal-Bench v2.1 at 57 and AA-LCR at 71. Priced at $0.30\u002F$1.20 per million input\u002Foutput tokens, the model is positioned for enterprise agent reliability.","South Korean AI company Upstage shipped Solar Pro 4, a closed commercial LLM, on August 11, 2026. The new generation is no longer pitched as \"bigger, stronger, better on general benchmarks\" — it is framed around \"agents that don't silently fail\", with retraining and evaluation rebuilt around long documents, terminal tasks, multi-turn tool use, and refusal-to-fabricate behaviors.\n\n## A block of agent-task scores\n\nSolar Pro 4 ships with a 512K context window and a 128K output ceiling, supports English, Korean, and Japanese, and reasons by default. Upstage's blog lists these agent scores from the Artificial Analysis evaluation system:\n\n- Terminal-Bench v2.1 (multi-step tasks in a live shell): 57\n- τ³-Banking (multi-turn tool use to find the right policy in a knowledge base): 23\n- AA-LCR (reasoning across roughly 100k tokens of long documents): 71\n\nUpstage's own side-by-side against Solar Open 2, on the same evaluation environment: Terminal-Bench v2.1 +13.8, BrowseComp +11.9, GDPval-AA v2 +7.4, AA-LCR +8.3. On GPQA Diamond, MMLU-Pro, LiveCodeBench, AIME 2026, the two models are roughly even. Artificial Analysis gives Solar Pro 4 a composite index of 42 — Upstage describes this as more than triple Solar Pro 3 and ahead of Nvidia's Nemotron 3 Ultra at 38 and Google's Gemini 3.5 Flash-Lite at 37.\n\nIn an August 20 piece in The New Stack, Upstage's US CEO Kasey Roh summarized the pitch: \"Save the frontier models for the frontier problems; we built the workhorse.\" Her concrete cost comparison: a document fact-checking task running 300K input plus 15K output tokens costs \"roughly $1 per task on premium frontier pricing versus $0.10 on Solar Pro 4.\"\n\n## Training core: synthesizing finished work with OfficeVerse\n\nSolar Pro 4 is trained on work synthesized by Upstage's in-house OfficeVerse pipeline. OfficeVerse generates office tasks from real public data across 11 industry domains and 12 task types, and grades each one pass-or-fail on the final deliverable. Upstage states that Solar Pro 4 was trained and validated on work in the same shape it takes in the real world. The same pipeline produced Ko-GDPval, Upstage's Korean office-work benchmark.\n\nThis is the philosophical opposite of benchmark-focused training. Many closed models top public leaderboards while still failing in production on documents: \"the model handles rows 1 through 200 just fine, then somewhere past that point it starts skimming — dropping rows, or silently filling a cell by pattern-matching from earlier rows instead of reading the actual value.\" Roh named this failure mode in The New Stack interview: \"That failure mode is precisely what we trained against.\"\n\n## \"Cannot verify\" is the headline deliverable of this generation\n\nSolar Pro 4 draws three explicit verdicts when answering document questions: grounded (the clause exists and gets a citation), not-in-document (the clause does not exist), mismatch (two documents differ on the same number — flagged as a discrepancy). When the evidence is missing, the model says \"cannot verify\" rather than producing a language-shaped answer from a learned prior.\n\nFor regulated industries — financial services, insurance, manufacturing, supply chain — this matters more than another leaderboard tick. Once a model invents a number or clause that isn't in the source document, it flows downstream unchanged, and eventually someone reconciles it manually. Solar Pro 4's design choice is to make traceability the default.\n\n## Pricing and availability\n\nSolar Pro 4 ships on Upstage Console, OpenRouter, Hermes Agent (Nous Research's agent environment), and Upstage Studio. Pricing: $0.30 per 1M input tokens, $0.06 per 1M cached input tokens, $1.20 per 1M output tokens. Upstage runs a launch promotion at 90% off list price from August 11 through September 10, 23:59 UTC.\n\nThe New Stack cites Upstage's figures: Solar Pro 4 crossed 370 billion tokens consumed on OpenRouter within a week of listing. Upstage's existing partners include AWS and AMD; the company opened a San Jose US headquarters in 2025.\n\n## A take\n\nSolar Pro 4 is not pitched as \"bigger than Sonnet or GPT\" or \"the next round of model capability arms race.\" It is a bet that, after several rounds of leaderboard saturation on SWE-Bench, Terminal-Bench, and AA-LCR, what actually decides whether an agent runs in production is whether the model can hold accuracy through hundreds of tabular rows, dozens of contract pages, and hundreds of tool calls without slipping into silent failure. That is an engineering problem, not a scaling problem. For developers everywhere — Chinese, Korean, US, European — Solar Pro 4 is a useful signal to watch: the next constraint in agent deployment is not parameter count but silent-failure rate.","upstage-solar-pro-4-agent-reliability-closed-llm","2026-08-25T03:00:00Z","2026-08-24T01:08:30.532075Z","2026-08-24T01:08:30.532095Z",true,"agent",54,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"d1e8997e-bb60-453d-9ef8-71b8bdde5386","Harvey 首个自研法律模型 Tenet 曝光:底座没选 GPT 和 Claude,选了 Kimi K3","harvey-tenet-kimi-k3-legal-model","2026-08-18T17:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00+00:00"]