[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-rosetta-maop":3,"topics-all":31,"news-related-386cd824-56cc-4474-8388-8caef2dbb9dd":50},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":24,"published_at":25,"created_at":26,"modified_at":27,"is_published":28,"publish_type":29,"image_url":13,"view_count":30},"386cd824-56cc-4474-8388-8caef2dbb9dd","Rosetta：让多模态预训练不再遗忘——腾讯混元与港科大提出 MAOP 零开销投影法","多模态大模型向新模态扩展时普遍被\"表征覆盖\"问题困扰：文生图等生成任务的高方差梯度会破坏已有语言能力，传统 MoE 路由容易 routing collapse，结构化 MoT 方案又切断跨模态协同。HKUST 与腾讯混元提出的 Rosetta 保留全局共享 QKV、FFN 解耦为可插拔任务专属专家 + Global Shared Expert，并提出 MAOP (Momentum-Anchored Orthogonal Projection)——复用 Adam 动量状态作为语义锚点、对新模态梯度做正交投影，全程零额外显存；Transfusion 框架下严格等活跃参数实验显示 Rosetta 同时在 MMLU 与文生图基准击败 MoE 与 MoT 基线，代码与权重已开源。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00293","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[],"tencent-rosetta-maop","2026-07-03T00:02:00Z","2026-07-03T00:08:43.223328Z","2026-08-19T02:08:40.142862Z",true,"agent",122,[32,41],{"slug":33,"tag_slug":33,"title_zh":34,"title_en":35,"intro_zh":36,"intro_en":37,"id":38,"is_active":28,"created_at":39,"modified_at":40},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":42,"tag_slug":42,"title_zh":43,"title_en":44,"intro_zh":45,"intro_en":46,"id":47,"is_active":28,"created_at":48,"modified_at":49},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":51},[52,57,62,67,72,77],{"id":53,"title":54,"news_slug":55,"published_at":56},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"70529122-522e-405d-9735-fc083706792f","数据重复 4 倍就开始退化:斯坦福UW团队实测 MoE 比稠密模型更怕数据墙","moe-overfit-repeated-data-stanford-uw","2026-09-13T19:20:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f5c0b227-faf9-47e3-863f-3c102365cd41","LongCat-Next 开源：把文字、图像和声音统一成离散 Token","longcat-next-discrete-native-multimodal","2026-08-09T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"d7b6d14d-7257-4794-b92f-31956bbc7eae","原生多模态 vs 后训练加压:国产头部基模两条路线的工程账","native-multimodal-vs-posttraining-2026","2026-08-05T00:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"2d29aa3d-317c-4126-a6f7-2c9c2c3b6f93","Kimi K3与DeepSeek V4之间,隔着原生多模态的时间差","kimi-k3-deepseek-v4-native-multimodal-divergence","2026-08-04T08:02:10+00:00"]