[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mixedbread-toast-1-search-subagent":3,"news-related-6953e0b7-8762-49fd-8b44-217944ebc0ea":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"6953e0b7-8762-49fd-8b44-217944ebc0ea","检索循环外包给专用小模型:Mixedbread Toast 1 官方称打平 Opus 5,token 少用 3.5 倍","Mixedbread 发布首款专用搜索智能体模型 Toast 1：它整体接管智能体工作流里的检索循环——自行拆解子查询、调用搜索工具、核查来源、整理证据包。官方称搜索质量打平或超过 Claude Opus 5 与 GPT-5.6 Sol，价格最高便宜 10 倍、速度快 12 倍；法律基准上同分情况下 token 少用 3.5 倍。所有关键数字均为厂商自报，第三方评测前宜观望。","\u003C!-- CURRENT TARGET: Mixedbread Toast 1, NOT any other model -->\n嵌入与重排模型的老玩家 Mixedbread 发布了第一款专用搜索智能体模型 Toast 1：它把智能体工作流里最烧钱的检索循环整体接管——自己拆子查询、调用搜索工具、读源、筛证据，只把整理好的上下文交回主模型。官方称其搜索质量打平或超过 Claude Opus 5 与 GPT-5.6 Sol，价格最高便宜 10 倍、速度快 12 倍（[官方发布](https:\u002F\u002Fwww.mixedbread.com\u002Fblog\u002Ftoast-1)）。\n\n## 逻辑：让贵的模型只干贵的活\n\nToast 1 的定位是「专用检索子代理」。它可以独立跑深度搜索，也可以作为 subagent 挂进 Codex 这类编码智能体——主模型负责推理与决策，取证环节整体外包。官方演示里，一条就业率比较查询由它拆成 16 次工具调用、3 轮完成，耗时 5.33 秒。上下文窗口 131K tokens，通过标准 Chat Completions API 接入，也提供 OpenCode 集成和开源 harness。\n\n## 两个关键成绩（均为厂商自报）\n\n**金融基准 OfficeQA Pro V2**（Databricks 发布，90 题）：GPT-5.6 Sol 在 Codex 中挂上 Toast 1 作子代理后，答案正确率达到 70%、单任务成本约 1.15 美元——官方称这是 Databricks 评测系统里的最高分；对照 Claude Fable 5（Databricks Genie）为 60% @ 约 4 美元，GPT-5.6 Sol（Codex）不挂 Toast 1 只有 33%。\n\n**法律基准 Harvey LAB 律所知识库**（随机 33 任务子集）：三种检索配置的任务得分完全相同（55 分），但 token 消耗从原生智能体的 80.6M 降到 23.0M（先换 Mixedbread Search 省 42%，再加 Toast 1 再省 51%），轮次从 21.7 降到 11.2——同分，token 少用 3.5 倍，官方称成本降幅超 60%。\n\n标准配置下单次搜索成本约 0.016–0.023 美元、中位延迟 8 秒；最高质量的 fusion 配置约 0.05–0.07 美元、11 秒。官方称在同等性能的系统里便宜 7–11 倍，而对照的前沿模型检索智能体延迟在 20 秒到 4 分钟之间。\n\n## 定价与边界\n\n发布期定价：输入 0.30 美元\u002F百万 tokens、缓存输入 0.036 美元（缓存写免费）、输出 0.72 美元\u002F百万 tokens，新用户送 5 美元额度。值得注意的边界：它是闭源 API 模型而非开放权重；BenchLM 明确将其标注为专有模型，且因官方自报成绩没有开放协议的评测表，暂不给排名——70% 那项成绩属于 GPT-5.6 Sol+Toast 1 组合系统，不是 Toast 1 单独跑分。官方也在脚注里承认，同类专用搜索智能体还有 SID-1 与 Chroma 的 Context-1，这个品类正在快速成型。\n\n## 我的看法\n\nToast 1 最值得记住的不是某个跑分，而是它把「智能体经济学」摆上了台面：当推理模型的单 token 价格不再降、而智能体任务又动辄百万 token 级检索时，把检索循环外包给一个便宜的专用小模型，是比“等模型降价”更现实的降本路径。但所有关键数字都来自厂商自报，第三方复现之前宜观望；它也再次印证了行业的分工趋势——通用推理归前沿大模型，脏活累活归专用子代理。对做 RAG 和研究型智能体的团队，5 美元额度值得一试。","https:\u002F\u002Fwww.mixedbread.com\u002Fblog\u002Ftoast-1","1c8024be-b3d7-4160-bf53-e352743c2fbe",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"045c011e-e2bb-45ce-bdd6-0c927f8a3b87","token-efficiency",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"b67f5595-7099-457a-a4ec-d3597c613d84","en","Mixedbread Toast 1: retrieval outsourced, 3.5x fewer tokens","Mixedbread has launched Toast 1, its first specialised search agent model. It takes over the entire retrieval loop of agentic workflows — decomposing queries, calling search tools, inspecting sources, and curating evidence. The company says search quality matches or outperforms Claude Opus 5 and GPT-5.6 Sol at up to 10x lower cost and 12x faster, with 3.5x fewer tokens at equal scores on a legal benchmark. All key numbers are vendor-reported; worth watching until independent replication arrives.","Mixedbread, a veteran maker of embedding and reranking models, has launched Toast 1, its first specialised search agent model. Toast 1 takes over the entire retrieval loop of agentic workflows — it decomposes queries into subqueries, calls search tools, inspects sources, and hands back only curated context. The company says its search quality matches or outperforms Claude Opus 5 and GPT-5.6 Sol while being up to 10× cheaper and 12× faster ([official announcement](https:\u002F\u002Fwww.mixedbread.com\u002Fblog\u002Ftoast-1)).\n\n## The logic: let expensive models do expensive work\n\nToast 1 is positioned as a dedicated retrieval subagent. It can run deep search standalone or plug into coding agents like Codex — the main model handles reasoning and decisions while evidence gathering is fully outsourced. In the official demo, an employment-rate comparison query was split into 16 tool calls across 3 rounds, resolved in 5.33 seconds. The context window is 131K tokens, accessible via a standard Chat Completions API, with an OpenCode integration and an open harness.\n\n## Two headline results (all vendor-reported)\n\n**OfficeQA Pro V2** (released by Databricks, 90 questions): GPT-5.6 Sol running in Codex with Toast 1 as a sub-agent reached 70% answer correctness at roughly 1.15 USD per task — per Mixedbread, the highest score among the systems Databricks evaluated. For comparison, Claude Fable 5 on Databricks Genie reached 60% at about 4 USD per task, while GPT-5.6 Sol in Codex without Toast 1 managed only 33%.\n\n**Harvey LAB's law-firm knowledge benchmark** (a randomly selected subset of 33 tasks): all three retrieval configurations scored an identical 55, but token consumption dropped from 80.6M (vanilla agent) to 47.0M after swapping in Mixedbread Search (−42%), and further to 23.0M with Toast 1 as the search subagent (−51%), with turns per task falling from 21.7 to 11.2 — same score, 3.5× fewer tokens, and a cost reduction the company puts at over 60%.\n\nA standard run costs roughly 0.016–0.023 USD per query with an eight-second median latency; the highest-quality fusion configuration costs about 0.05–0.07 USD at eleven seconds. Among systems reaching similar performance, Mixedbread claims Toast 1 is 7–11× cheaper, while the frontier-model retrieval agents it compared against took 20 seconds to four minutes on the same evaluation.\n\n## Pricing and caveats\n\nLaunch pricing: 0.30 USD per million input tokens, 0.036 USD per million cached input tokens (cache writes free), and 0.72 USD per million output tokens, with 5 USD in credits for new users. Worth noting on the boundaries: this is a proprietary API model, not open weights; BenchLM explicitly lists it as proprietary and leaves it unranked because the self-reported results lack an open protocol benchmark table — the 70% figure belongs to the GPT-5.6 Sol + Toast 1 combination system, not Toast 1 standalone. Mixedbread's own footnote acknowledges sibling efforts like SID-1 and Chroma's Context-1: this category is taking shape fast.\n\n## My take\n\nThe most memorable thing about Toast 1 is not any single benchmark but the way it puts agent economics on the table. When per-token prices of reasoning models stop falling while agentic tasks routinely burn millions of retrieval tokens, outsourcing the retrieval loop to a cheap specialised model is a more realistic cost path than waiting for frontier prices to drop. That said, every key number here is vendor-reported — worth watching until independent replication arrives. It also confirms the industry's division of labour: general reasoning belongs to frontier models, and the grunt work goes to specialised subagents. For teams building RAG and research agents, the 5-USD credit is worth a try.","mixedbread-toast-1-search-subagent","2026-08-16T21:10:00Z","2026-08-16T21:07:02.529539Z","2026-08-16T21:07:02.529548Z",true,"agent",65,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"e75069c6-f15c-4ff9-8b11-404d705442e8","Upstage Solar Pro 4:把「agent 跑得稳」做成新一代闭源模型卖点","upstage-solar-pro-4-agent-reliability-closed-llm","2026-08-25T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"d1e8997e-bb60-453d-9ef8-71b8bdde5386","Harvey 首个自研法律模型 Tenet 曝光:底座没选 GPT 和 Claude,选了 Kimi K3","harvey-tenet-kimi-k3-legal-model","2026-08-18T17:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00"]