[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cat-confidence-adaptive-thinking":3,"topics-all":36,"news-related-fbf3ec38-2aba-4199-8e55-c56071ea6e24":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"fbf3ec38-2aba-4199-8e55-c56071ea6e24","CAT 让 LRM 不再「想太多」:把模型自我置信度变成推理长度调速器","DeepSeek-R1 之后业界发现反直觉事实:LRM 在「1+1=?」上也会写几百 token 反思,token 与延迟被严重浪费。arXiv 2607.00862 收录、被 ACL 2026 Industry Track 接收的 CAT(Confidence-Adaptive Thinking),把模型「自我置信度」当调速器,让 LRM 自己决定每道题该想多久。\n\n过去做推理压缩,要么「一刀切」压缩 CoT,要么外挂分类器粗判难度——前者难题上掉精度,后者易误判。CAT 的观察很简洁:模型自己其实「知道」每道题有多少把握,这种 self-certainty 在偏好优化里几乎没人用过。作者把 confidence 作为偏好信号直接灌进偏好优化,让模型同时学两件事:自信的题压缩回答,不自信的题充分推演。同一组权重,在「9.9 比 9.11」上两三行答完,在证明题上认真推几十步。多个 benchmark 稳定超过 SOTA,平均 token 消耗显著降低。\n\nCAT 的真正贡献不是新算法,而是指出一个被忽视的免费信号——模型对自身输出的置信度天然随题而变。当工业界把 RL 目标改成「长度自适应用户需求」时,这种「内省式」信号比外挂分类器更鲁棒、更便宜,后续在 agent 规划、多轮推理、工具调用场景里会进一步扩散。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00862","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c1e7aaac-21f1-4de4-977c-74314c8d7278","en","CAT turns self-confidence into a reasoning-length throttle","After DeepSeek-R1, the industry discovered a counter-intuitive fact: LRM writes hundreds of tokens of reflection even on \"1+1=?\", with token and latency being seriously wasted. CAT (Confidence-Adaptive Thinking), included in arXiv 2607.00862 and accepted by ACL 2026 Industry Track, treats the model's \"self-confidence\" as a governor, letting the LRM decide for itself how long to think on each problem. Previously, reasoning compression was done either by \"blanket\" CoT compression, or by attaching an external classifier to roughly judge difficulty — the former loses accuracy on hard problems, the latter is easily misjudged. CAT's observation is simple: the model actually \"knows\" how confident it is on each problem, and this self-certainty has rarely been used in preference optimization. The authors pour confidence as a preference signal directly into preference optimization, letting the model simultaneously learn two things: compress answers on confident problems, fully reason on unconfident problems. The same set of weights, two or three lines on \"9.9 vs 9.11\", carefully reason dozens of steps on proof problems. Stable SOTA on multiple benchmarks, with significant reduction in average token consumption. CAT's real contribution isn't a new algorithm, but pointing out a neglected free signal — the model's confidence in its own output naturally varies with the problem. When the industry changes RL objectives to \"length-adaptive user needs\", this \"introspective\" signal is more robust and cheaper than an external classifier, and will further spread in agent planning, multi-turn reasoning, and tool-calling scenarios.","cat-confidence-adaptive-thinking","2026-07-05T16:05:00Z","2026-07-05T16:08:09.536008Z","2026-08-19T02:08:40.142862Z",true,"agent",136,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"e648a701-6b9d-4a4b-9122-d2a4afc8349b","后训练方法终于有了统一坐标系：SFT、RLHF、Distillation 到底在做什么？","post-training-unified-coord-sft-rlhf-distill","2026-05-30T05:12:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"b571067a-9fa8-42bf-9431-98f26ac78e03","伯克利把LLM推理搬进SSD:KV缓存压缩15倍","llm-inference-in-flash-cim-ssd","2026-09-19T21:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b667e52f-ec7d-4ca4-8d9e-1db81e1a5616","DeepSeek论文:890字节KV缓存的三层架构账","deepseek-v41-flash-kv-cache-paper","2026-09-18T15:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"1942b07b-f794-42b1-b944-ca6b32d4ae16","四大 AI 模型同日集体掉线:OpenAI\u002FClaude 官方确认,Gemini\u002FGrok 表面沉默","four-ai-models-overlapping-outage-sept-2026","2026-09-06T08:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"31f3215c-0892-419d-a610-fe815cc60bbe","GPT-5.6 降价 80% 把竞争拉进「同等智能成本」：DeepSeek V4 Flash 接招，国产模型卡出双线赛道","gpt-5-6-luna-price-cut-equal-intelligence-cost","2026-08-12T03:00:00+00:00"]