[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemma-4-12b-encoder-free-multimodal-16gb":3,"topics-all":36,"news-related-208b4b5c-b2fe-4789-a928-f01de7a271b0":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"208b4b5c-b2fe-4789-a928-f01de7a271b0","Gemma 4 12B 发布：Google 开源多模态模型首次实现无编码器架构","Google DeepMind 于 2025 年 6 月 3 日发布了 Gemma 4 12B，这是一款参数规模为 120 亿的开源多模态模型，最大的亮点在于其采用了**无独立视觉\u002F音频编码器的架构设计**——所有模态直接流入同一个解码器 Transformer，视觉和音频信号通过轻量嵌入模块直接注入 LLM 主干网络，不再需要独立的编码器来处理图像和音频输入。这一设计使得模型体积大幅缩小，同时保留了强大的多模态理解能力。Gemma 4 12B 支持文本、图像、视频和原生音频的统一处理，能够理解视觉内容、处理音频输入并执行复杂推理任务。由于参数精度的优化，该模型可以在配备 16GB 显存的笔记本电脑上本地运行，满足边缘 AI 场景的需求。此外，它采用 Apache 2.0 许可证，对商业使用限制较少，适合开发者部署本地化 Agent 工作流。与 Google 此前发布的 Gemma 4 26B MoE 版本相比，12B 虽然参数更少，但在大多数标准 benchmark 上性能接近 26B，却只占用不到一半的显存。对于需要在本地设备上构建多模态 AI 能力的开发者来说，这是一款值得关注的新选择。","https:\u002F\u002Fdevelopers.googleblog.com\u002Fbringing-gemma-4-12b-to-your-laptop-unlocking-local-agentic-workflows-with-google-ai-edge\u002F","35ce748f-48b7-4638-88ef-effa57a7e749",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1aad9a73-bb7d-4fb3-ab58-214ad12835e1","en","Gemma 4 12B: Google's first encoder-free multimodal open model","Google DeepMind released Gemma 4 12B on June 4, the first open-source multimodal model with an encoder-free architecture. The \"encoder-free\" design skips the separate vision encoder, instead having the language model directly process image tokens. This simplifies the architecture, reduces latency, and enables more flexible multimodal interactions, with the model running on-device via Google AI Edge.","gemma-4-12b-encoder-free-multimodal-16gb","2026-06-04T10:05:00Z","2026-06-04T10:03:18.404680Z","2026-08-19T02:08:40.142862Z",true,"agent",150,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"1e2312f1-07dd-46d4-823c-1bb5cf620ed5","Gemma 4 技术报告:Google 把多模态推进到「无编码器原生」时代","gemma-4-tech-report","2026-07-08T08:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"039ff515-68e7-4f11-866a-1da97e26eb45","Gemini 3.8 Live 拿下 S2S 实时语音榜第一","gemini-3-8-live-voice-s2s-number-one","2026-09-15T17:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"089f56f3-32ff-4036-89b5-728d5f5a9359","边聊边干活:腾讯混元开源全模态交互 Agent Gander,小脑管对话、大脑管执行","hunyuan-gander-omni-interaction-agent","2026-09-09T21:07:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00"]