[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-coat-continuous-audio-thinking-latent":3,"topics-all":36,"news-related-0b9b2439-b5f8-413e-8e62-58621123cc2f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0b9b2439-b5f8-413e-8e62-58621123cc2f","Continuous Audio Thinking：把「思考」搬进音频 LLM，零解码成本补齐声学信息损失","大模型这条线最近一年最显著的方向就是「把思考做厚」——Chain-of-Thought、潜在 CoT、thinking tokens——但所有这些设计默认在文本空间里展开。CoAT (Continuous Audio Thinking) 这篇来自韩国研究团队的工作把视角切到了更上游的「声学」维度，指出当前 LALMs 的结构性痛点：因为训练目标只对齐文本响应，模型的隐藏状态会逐步向「文本友好」形态塌缩，原本承载音素细节、韵律、声学事件、情感、pitch 等关键信息的中间表征在层层传播中被稀释，到解码阶段早已不可用。\n\nCoAT 的做法相当克制——不引入任何新的自回归 token，而是让模型在 prefill 阶段就把声学信息打点到一段连续潜在空间里，由音频专家模型蒸馏出辅助监督信号。关键工程优势是这段连续 thinking block 是一次性 prefill 处理的，不会在 inference 端增加额外 decode 成本。\n\n实验覆盖 Qwen2-Audio、Qwen2.5-Omni-7B、Audio Flamingo 3 三个不同架构的 LALM，在音频推理、音频理解、音乐分类、语音情感、语音转写五大类基准上同步取得提升。这种 plug-in 式设计意味着现有 LALM 不需要重新训练主线权重，换上一个 CoAT 模块即可获得跨任务增益。\n\n最有意思的是 paper 末尾对 thinking 位置辅助监督信号的传播分析——它证明声学信息并非被解码 token 真正消费，而是通过隐藏状态扩散到最终的文本响应中。这给「thinking trace 不需要显式 token」提供了新的实验证据，也意味着未来 LALM 可以用更低带宽、更经济的 latent thinking 路径替代纯文本 CoT。\n\n对国内做多模态语音模型的团队来说，CoAT 的方法论可以直接借鉴：用轻量 expert distillation head 把声学知识钉在中间层，比硬上 autoregressive thinking tokens 务实得多。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.18273","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"43585318-60c3-4f48-8506-076ecec08d39","en","Continuous Audio Thinking: reasoning inside audio LLMs","arXiv 2606.18273 introduces Continuous Audio Thinking (CAT), a method for bringing the \"thinking\" capability of reasoning models into audio LLMs. The result: audio LLMs can now \"think\" about audio input before responding, with no extra decoding cost.\n\nThe \"audio LLM thinking\" challenge: text-based reasoning models (e.g., OpenAI o1) can \"think\" by generating internal reasoning tokens before the final answer. This \"thinking\" is what enables complex multi-step reasoning. Audio LLMs traditionally don't have this capability — they go directly from audio input to text output, losing the \"thinking\" step.\n\nThe CAT fix: CAT adds a \"continuous thinking\" stage to the audio LLM, where the model generates an internal \"thought\" representation (a sequence of continuous vectors) before generating the final text response. The \"thought\" is generated in the audio LLM's continuous latent space, so there's no decoding cost — the thought is internal, not external.\n\nThe benchmark: CAT-augmented audio LLMs match text-based reasoning models on tasks that require understanding complex audio (multi-speaker conversations, music analysis, environmental sound reasoning). The biggest improvement is on \"audio QA with reasoning\" tasks, where CAT hits 15-20 point improvements over the baseline.\n\nThe bigger takeaway: \"thinking in latent space\" is a significant new direction. The \"thinking\" capability of reasoning models is currently limited to text — the model generates text tokens as the thought. CAT shows that \"thinking\" can be done in continuous latent space, with no decoding cost, and the technique generalizes to audio, image, and video. For the industry, this means \"latent thinking\" will become a standard feature of multimodal reasoning models.","coat-continuous-audio-thinking-latent","2026-06-18T06:00:00Z","2026-06-18T06:11:16.793693Z","2026-08-19T02:08:40.142862Z",true,"agent",178,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"efca492c-b9b1-4d72-bf3b-32f08fc0f515","世界模型崛起：AI 从数字世界走向物理世界的关键一步","mit-world-models-physical-ai-lecun-fei-fei","2026-04-27T13:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"d17a841b-abca-46e0-80e4-d955f1c837ba","亚马逊 VGT3 仓库曝光:一天拆掉上千本书,只为给 AI 模型喂语料","amazon-vgt3-warehouse-ai-training-books","2026-09-07T03:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"1464179a-2b7f-4369-b680-25868ddd9042","皮尤实测：超过三分之一 ChatGPT 后的英文网页已有 AI 写作痕迹","pew-research-ai-web-content-2026","2026-08-31T03:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"21a91da5-c5fa-45e2-b01f-a7950331cf44","S3 把 DuckDB 团队收走了:DuckLabs 加盟 AWS,MIT 开源照旧","aws-buys-ducklabs-duckdb-open-source","2026-08-30T06:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"216f3c2b-d551-45fc-a206-c3ccfae9db89","亚马逊 Mechanical Turk 将永久关闭:被 AI 掏空的众包平台","amazon-mechanical-turk-shutdown","2026-08-29T17:30:00+00:00"]