[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemma-4-tech-report":3,"topics-all":36,"news-related-1e2312f1-07dd-46d4-823c-1bb5cf620ed5":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1e2312f1-07dd-46d4-823c-1bb5cf620ed5","Gemma 4 技术报告:Google 把多模态推进到「无编码器原生」时代","7 月 2 日挂上 arXiv 的 Gemma 4 技术报告,Gemma Team 300 余位作者合力完成。整套体系覆盖 2.3B 到 31B 共五档,**首次在主版本里同时给出 dense 与 MoE 两条路线**:小尺寸守「手机可跑」,大尺寸用 MoE 拉容量、压激活。\n\n12B 最值得展开。Gemma 4 把视觉与音频编码器彻底拿掉,改成 encoder-free 结构,直接吞 raw 图像 patch 与音频波形,把多模态融合从「外挂器官」变成「原生器官」。这与 Qwen2.5-Omni、LLaVA 路线截然不同;代价是训练稳定性更难,Google 把它放在 12B 而非旗舰尺寸,先把成本压下来再往上推。\n\nThinking mode 落地,先思考 token 再回答,做成可开关能力。配合长上下文优化,Gemma 4 在 STEM、多模态、长上下文 benchmark 上明显跃升,在 human-rated 任务上对位更大的开源前沿模型。\n\n把 Google 在 Gemini 闭源体系里验证过的多模态、长上下文与推理范式,**用开源可复现的方式重新写一遍**——这才是这份报告真正的分量。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.02770","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8601a63f-87ea-4181-9273-aa5e067a2de0","en","Gemma 4 tech report: the encoder-free multimodal era","The Gemma 4 tech report, posted to arXiv on July 2, was completed by over 300 authors from the Gemma Team. The system covers five scales from 2.3B to 31B, **for the first time in the main version simultaneously providing both dense and MoE paths**: the small sizes keep \"runnable on phones\", while the large sizes use MoE to expand capacity and compress activation. The 12B is most worth elaborating. Gemma 4 completely removes the visual and audio encoders, switching to an encoder-free structure, directly swallowing raw image patches and audio waveforms, making multimodal fusion from \"external organ\" to \"native organ\". This is completely different from the Qwen2.5-Omni, LLaVA route; the cost is harder training stability, and Google puts it at 12B rather than the flagship size, first getting the cost down before pushing it up. Thinking mode lands, first think tokens then answer, made into a switchable capability. Combined with long-context optimization, Gemma 4 makes a clear jump on STEM, multimodal, and long-context benchmarks, and on human-rated tasks it stands opposite larger open-source frontier models. The point of taking the multimodal, long-context, and reasoning paradigms validated by Google in the closed-source Gemini system and **rewriting them in an open-source reproducible way** — that's the real weight of this report.","gemma-4-tech-report","2026-07-08T08:05:00Z","2026-07-08T16:13:48.356739Z","2026-08-19T02:08:40.142862Z",true,"agent",161,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"208b4b5c-b2fe-4789-a928-f01de7a271b0","Gemma 4 12B 发布：Google 开源多模态模型首次实现无编码器架构","gemma-4-12b-encoder-free-multimodal-16gb","2026-06-04T10:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"039ff515-68e7-4f11-866a-1da97e26eb45","Gemini 3.8 Live 拿下 S2S 实时语音榜第一","gemini-3-8-live-voice-s2s-number-one","2026-09-15T17:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"089f56f3-32ff-4036-89b5-728d5f5a9359","边聊边干活:腾讯混元开源全模态交互 Agent Gander,小脑管对话、大脑管执行","hunyuan-gander-omni-interaction-agent","2026-09-09T21:07:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00"]