[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-diffusiongemma-26b-google-4x-speedup":3,"topics-all":36,"news-related-b4f3ab96-75a7-4565-9a6f-cd21419928d8":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b4f3ab96-75a7-4565-9a6f-cd21419928d8","DiffusionGemma 26B 开源：Google 把扩散范式搬进文本生成，单卡 4× 加速 Gemma 4","Google 在 6 月 10 日的 AI Blog 推出 DiffusionGemma——一个 26B 总参数 \u002F 3.8B 激活的 MoE 实验模型，Apache 2.0 许可开源。它不是 Gemma 4 家族的常规迭代，而是一次范式跃迁：把图像\u002F视频领域的扩散机制搬进文本生成，单次前向 256 token 并行解码，宣称在 H100 上达到 1000+ t\u002Fs、RTX 5090 上 700+ t\u002Fs，相对 Gemma 4 自回归版最多 4 倍加速。\n\n**技术核心**。DiffusionGemma 走的是「占位 token 画布 + 多次迭代去噪」路线，与自回归「从左到右」逐 token 推理完全不同。所有 token 在生成时都可 attend 到整段文本——对在线编辑、代码 infill、数学图、氨基酸序列等非线性任务天然友好。MoE 设计 + 量化让它能塞进 18GB 显存，单张 RTX 4090\u002F5090 即可跑。代价是整体输出质量低于标准 Gemma 4，Google 明确把它定位为「实验」和「为本地、低并发、交互场景而生」。\n\n**生态协同**。Google 同步给了一整套工程栈：Hugging Face、vLLM（Red Hat 集成）、MLX、Unsloth、NVIDIA NeMo 全部 day-0 支持，llama.cpp 即将到来。原生 NVFP4 4-bit 浮点让 Hopper\u002FBlackwell 上的吞吐再上一个台阶。Unsloth 已经放出 Sudoku 和 3D SVG 微调 demo。\n\n**评论**。这条路线延续了 Gemini Diffusion 的研究脉络，但更值得关注的不是它会不会替代 GPT\u002FGemini 主线，而是它和 Nemotron-Labs Diffusion、DFlash 一起，构成 2026 年「扩散语言模型从论文走向可用工具」的拐点。在云端高 QPS 场景下，自回归模型靠 batching 仍占优；但对个人开发者、笔记本玩家、单卡工作站来说，4× speedup 是实打实的体验质变。文本扩散也许不会成为主流，但会成为边缘侧、实时编码、本地 IDE 的强力补充。","https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Ftechnology\u002Fdevelopers-tools\u002Fdiffusion-gemma-faster-text-generation\u002F","4d11edad-2df6-45f6-b71f-70f65de7f7fd",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b00039f3-2efc-4d57-bf07-02a6f29e02ff","en","DiffusionGemma 26B: diffusion text generation, 4x on one GPU","On June 10 Google AI Blog launched DiffusionGemma — a 26B-total \u002F 3.8B-active MoE experimental model, open-sourced under Apache 2.0. It is not a routine iteration of the Gemma 4 family, but a paradigm leap: the diffusion mechanism from image\u002Fvideo is brought into text generation, decoding 256 tokens in parallel in a single forward pass, claiming 1000+ t\u002Fs on H100 and 700+ t\u002Fs on RTX 5090, a maximum 4× speedup over the Gemma 4 autoregressive version.\n\n**Technical core.** DiffusionGemma takes a \"placeholder-token canvas + multi-iteration denoising\" path, fundamentally different from autoregressive \"left-to-right\" token-by-token inference. All tokens during generation can attend to the full text — naturally friendly to non-linear tasks like online editing, code infill, math graphs, and amino acid sequences. The MoE design + quantization lets it fit in 18GB of VRAM, runnable on a single RTX 4090\u002F5090. The cost is that overall output quality is lower than standard Gemma 4, and Google explicitly positions it as \"experimental\" and \"born for local, low-concurrency, interactive scenarios.\"\n\n**Ecosystem synergy.** Google shipped a full engineering stack at the same time: Hugging Face, vLLM (Red Hat integration), MLX, Unsloth, NVIDIA NeMo all support it on day-0, with llama.cpp coming soon. Native NVFP4 4-bit floating point gives Hopper\u002FBlackwell an additional throughput bump. Unsloth has already released Sudoku and 3D SVG fine-tuning demos.\n\n**Commentary.** This path continues the research line of Gemini Diffusion, but the more notable thing isn't whether it will replace the GPT\u002FGemini main line, but that together with Nemotron-Labs Diffusion and DFlash it forms the 2026 inflection point where \"diffusion language models go from papers to usable tools.\" In high-QPS cloud scenarios, autoregressive models still win by batching; but for individual developers, laptop users, and single-GPU workstation players, 4× speedup is a real qualitative change. Text diffusion may not become mainstream, but it will become a powerful complement for edge, real-time coding, and local IDE scenarios.","diffusiongemma-26b-google-4x-speedup","2026-06-10T14:00:00Z","2026-06-10T16:10:47.349071Z","2026-08-19T02:08:40.142862Z",true,"agent",129,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4455b8ee-eab9-463b-9934-f1df4b1b4fb3","扩散语言模型的适配断点被接上:dQwen3.5 只花一半 token","dqwen3-5-hybrid-attention-diffusion-language-models","2026-09-18T19:20:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"20568e5d-3b66-495d-8b7b-0a702f3c7877","模型在进化,训练环境却是死的:Google 开源 EnvHarness,给环境也套一层 harness","google-envharness-agent-environments","2026-08-20T10:42:06+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"1e2312f1-07dd-46d4-823c-1bb5cf620ed5","Gemma 4 技术报告:Google 把多模态推进到「无编码器原生」时代","gemma-4-tech-report","2026-07-08T08:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"bce625bf-8835-47d1-a12e-bf0cc111b905","dOPSD：让扩散 LLM 用「自身去噪轨迹」当老师，Dream 与 LLaDA 数学、代码双涨","dopsd-diffusion-llm-self-distillation","2026-07-07T06:05:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"9d376eae-46cf-48df-acd5-f19994948428","SLIM-RL:扩散大模型 RL 训练从「轨迹重构」走向「风险控制」","slim-rl-diffusion-rl-risk","2026-07-05T20:30:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"1fe65909-67b4-4019-89fe-76eb4b226c41","Interfaze 把扩散解码塞进 ASR：diffusion-gemma-asr-small 用 42M 适配器撬开六语种开放语音识别","interfaze-diffusion-gemma-asr","2026-07-04T04:10:00+00:00"]