[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-claude-code-prompt-cache-cost":3,"topics-all":36,"news-related-be171a5b-9d4a-48d7-8660-ecda8abb85f7":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"be171a5b-9d4a-48d7-8660-ecda8abb85f7","Claude Code 上 2908 次对照实验撕开「压缩省钱」幻觉:prompt cache 才是账单大头","一个常被忽视的事实正在重塑 API 型 coding agent 的优化思路:cache creation 与 cache reads 在 Claude Code 账单中占比高达 87%(按四分量重建口径)或 80%(按实际账单口径)。Sarel Weinberger 与 Amir Hozez 在 arXiv:2607.12161 中用 hash-frozen 的 2,908 次 Claude Code 实测(2,848 次纳入分析、103 个任务、7 个仓库、3 个模型)直接否定了「压缩工具输出=省钱」的直觉。其中一组 arm 砍掉 38% 工具输出 token 后,配对成本反而上升 6.8%(95% CI: +2.8% 到 +11.3%),任务级相关系数仅 0.15,置信区间跨零。更危险的是压缩对动作证据的破坏:SWE-bench 派生 Go 任务上,压缩后 patch 成功率从 27\u002F40 跌到 15\u002F40,源于字面编辑锚点被改坏。论文据此提出 success-adjusted billed cost 作为新基准,提示业界把优化焦点从「压缩率」转向「缓存命中结构与证据保真」。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.12161","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"337c6827-9a18-46c9-9ba3-09038ee20194","en","2,908 Claude Code trials: prompt cache drives the real bill","A frequently overlooked fact is reshaping the optimization thinking for API-style coding agents: cache creation and cache reads account for up to 87% of the Claude Code bill (by a four-component reconstruction view) or 80% (by actual billing). Sarel Weinberger and Amir Hozez, in arXiv:2607.12161, use 2,908 hash-frozen real Claude Code experiments (2,848 analyzed, 103 tasks, 7 repos, 3 models) to directly refute the intuition that \"compressing tool output = saving money\". In one arm, after cutting 38% of tool-output tokens, the paired cost actually rose 6.8% (95% CI: +2.8% to +11.3%), the task-level correlation coefficient was only 0.15, and the confidence interval crossed zero. More dangerously, compression destroys the action-evidence chain: on SWE-bench-derived Go tasks, post-compression patch success rate dropped from 27\u002F40 to 15\u002F40, because literal edit anchors got mangled. The paper accordingly proposes success-adjusted billed cost as a new baseline, urging the industry to shift the optimization focus from \"compression ratio\" to \"cache hit structure and evidence fidelity\".","claude-code-prompt-cache-cost","2026-07-15T18:35:00Z","2026-07-15T18:15:13.586997Z","2026-08-19T02:08:40.142862Z",true,"agent",129,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4bbc55d2-cabc-477f-a3ad-4e2c119aff2a","TokTier 抓住 Agent 推理的隐藏瓶颈：缓存命中 94.1%，分词仍吃掉 64% 首 token 时间","toktier-stateful-tokenization-agent-serving","2026-07-31T17:56:30+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"6b52b4a9-d567-46b8-99c1-e9c65ba59b16","SWE-Pruner Pro:ByteDance 让 Agent 自己当剪枝器,省 39% token 还涨分","swe-pruner-pro-bytedance","2026-07-25T12:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"37aa0bc9-d135-444f-842e-0b40388d29e9","Qwen3.7-Max 原生兼容 Anthropic API 协议：Claude Code 现已可直接调用阿里模型","qwen3-7-max-anthropic-api-claude-code","2026-05-27T10:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"57b691cc-476c-4427-8618-e29127654b34","AMD ROCm 7 原生支持 Qwen3-Coder-Next：单卡 256k 上下文打破推理硬件垄断","amd-rocm7-qwen3-coder-next-256k-mono","2026-05-25T16:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"0d0e5ce8-fa18-4907-b811-2918ff8464e4","FlexSQL：小型LLM如何在Text-to-SQL任务上超越GPT-o3和DeepSeek-R1","flexsql-nus-text-to-sql-spider2-65pct-gpt-oss-120b","2026-05-05T10:15:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"69e52a42-19c7-4580-8c49-5446233fbdde","7B模型如何超越GPT-4o？ICLR Oral论文揭示AgentFlow流式训练新范式","agentflow-7b-icrl-oral-flow-grpo-14-9pct","2026-05-03T01:10:00+00:00"]