[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-apple-pararnn-7b-rnn-665x-faster-iclr":3,"topics-all":36,"news-related-2ca620f4-a046-4274-925d-f0689123a498":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"2ca620f4-a046-4274-925d-f0689123a498","ParaRNN：Apple 让 RNN 重回战场，7B 参数模型训练提速 665 倍","RNN（循环神经网络）在 LLM 领域几乎被 Transformer 完全取代，但 Apple 最新研究正在改变这一格局。Apple 机器学习研究团队在 ICLR 2026 发表论文 ParaRNN，提出一种并行化训练框架，首次让 RNN 能够在数十亿参数规模上进行高效训练。在实验中，ParaRNN 将传统顺序训练的 RNN 速度提升 665 倍，成功训练出首个 70 亿参数、性能与 Transformer 相当的 RNN 模型。RNN 天然比 Transformer 更节省内存和计算资源，推理效率优势明显。但其核心瓶颈在于——时间步必须顺序计算，无法并行。这使得大规模 RNN 训练成本极高，长期被学术界和工业界搁置。ParaRNN 通过重新设计 RNN 的计算图，实现了训练时的完全并行化。团队将改造后的 GRU 和 LSTM 单元（ParaGRU、ParaLSTM）应用于大规模语言建模任务。结果显示，在相同参数规模下，ParaRNN 训练的 RNN perplexity 与 Transformer 和 Mamba2 相当，甚至略有优势。这项突破的实际意义在于：它为未来 LLM 架构选择提供了新的可能性。当部署场景对推理延迟和内存高度敏感时，RNN 作为一种低资源方案，重新进入了可选项。代码已公开释放。","https:\u002F\u002Fmachinelearning.apple.com\u002Fresearch\u002Fpararnn","a2e6145a-2a88-4c51-8d09-c4375b2a833b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8ef0bbdb-e9f4-4b9a-aa1b-8bc399dcd5e5","en","ParaRNN: Apple brings RNNs back, 665x faster training at 7B","Apple Machine Learning Research released ParaRNN, a parallel RNN architecture that brings the recurrent neural network back as a serious contender for sequence modeling. The standout: a 7B-parameter RNN model trains 665× faster than equivalent Transformer models, with competitive quality on language tasks. The \"RNN comeback\" could shift the long-context inference landscape.","apple-pararnn-7b-rnn-665x-faster-iclr","2026-05-27T13:15:00Z","2026-05-27T13:15:52.291724Z","2026-08-19T02:08:40.142862Z",true,"agent",216,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f2b0bde3-ebb5-49e5-a418-5ece37639d1b","MIPU\u002FMIPI：把 LLM RL 的「训练—推理失配」从工程噪音重写为优化目标","mipu-mipi-rl-mismatch","2026-07-04T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"5a4c2a98-6ce7-4e51-8348-118be3083afc","ELDR 把 MoE 推理的「延迟最后一公里」拉直:vLLM 实测 TPOT 最多砍 13.9%","eldr-moe-routing","2026-07-03T12:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"cd82f869-d4a7-45b7-95ce-b66051e9d933","BlockPilot：实例自适应策略学习让扩散式投机解码再下一城,Qwen3-4B 上首破 4.20× 加速","blockpilot-instance-adaptive-block","2026-07-01T14:17:29+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"abf22bbb-eccf-46d0-99d6-debe1596f92b","自验证蒸馏：无需外部教师，LLM如何实现自我进化","self-verified-distill-qwen3-16-7pp-math","2026-05-27T19:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"6c9e553d-3029-4867-bc29-ee069d26b934","自验证蒸馏：无需外部教师，LLM 如何实现自我进化","svd-self-verified-distill-stanford-perplexity","2026-05-27T16:05:00+00:00"]