[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-sana-wm-2-6b-world-model-720p":3,"topics-all":36,"news-related-aa88f4b4-bb5c-4411-9ab3-18929dbd4444":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"aa88f4b4-bb5c-4411-9ab3-18929dbd4444","SANA-WM：NVIDIA 26 亿参数开源世界模型，单卡分钟级 720p 视频生成","NVIDIA 近日发布了 SANA-WM，一个 26 亿参数的开源世界模型，能够在单张 GPU 上生成一分钟长度、720p 分辨率的视频，并支持公制尺度 6-DoF 相机控制。该模型已在 arXiv 公开（arXiv:2605.15178），代码和权重均可在 NVlabs\u002FSana GitHub 仓库获取。\n\n架构核心是 Hybrid Linear Attention：用帧级 Gated DeltaNet（GDN）替代大部分 Attention 块，引入衰减门 γ 解决长视频状态漂移问题，让 recurrent state 保持在常量维度。\n\n两阶段 pipeline：第一阶段生成低分辨率粗略输出，第二阶段通过长视频精修器提升质量。经 4 步蒸馏的版本在单张 RTX 5090（NVFP4 量化）上完成 60 秒 720p 视频去噪仅需 34 秒，吞吐量是此前开源方案的 36 倍。训练仅需约 21.3 万段公开视频，在 64 张 H100 上训练 15 天即可完成。\n\nSANA-WM 的意义在于让世界模型的训练和推理都能在有限算力下完成。当分钟级、720p、带相机控制的视频生成可以在消费级硬件上运行，世界模型作为机器人仿真和具身智能训练数据来源的实用价值才真正打开。从「能跑」到「用得起」，这是 2026 年世界模型领域最务实的一步。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.15178","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ce77cdc7-5f7f-419e-885f-e069edccacac","en","SANA-WM: NVIDIA's 2.6B world model, minute-scale 720p video","NVIDIA recently released SANA-WM, a 2.6B-parameter open-source world model capable of generating one-minute-long 720p-resolution video on a single GPU, with metric-scale 6-DoF camera control. The model is publicly available on arXiv (arXiv:2605.15178), and code and weights are accessible in the NVlabs\u002FSana GitHub repository.\n\nThe architectural core is Hybrid Linear Attention: most attention blocks are replaced by frame-level Gated DeltaNet (GDN), introducing a decay gate γ to address long-video state drift, keeping the recurrent state at constant dimension.\n\nTwo-stage pipeline: the first stage generates low-resolution rough output, and the second stage improves quality through a long-video refiner. The 4-step distilled version completes 60 seconds of 720p video denoising in just 34 seconds on a single RTX 5090 (NVFP4 quantized), with throughput 36× that of previous open-source solutions. Training requires only about 213,000 public video segments, completed in 15 days on 64 H100s.\n\nSANA-WM's significance lies in making both the training and inference of world models achievable on limited compute. When minute-level, 720p, camera-controlled video generation can run on consumer-grade hardware, the practical value of world models as a data source for robot simulation and embodied-intelligence training truly opens up. From \"can run\" to \"affordable to use,\" this is the most pragmatic step for the world-model field in 2026.","nvidia-sana-wm-2-6b-world-model-720p","2026-05-17T02:06:00Z","2026-05-17T10:08:23.761165Z","2026-08-19T02:08:40.142862Z",true,"agent",180,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"f7234b7e-a2c3-404b-9fc4-aaca8e0c8f91","Edge0 预测路由:35B MoE 挤进 24GB Mac","edge0-prerouter-ssd-moe","2026-09-17T15:10:05+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"ff6f65e1-28b2-4a48-b317-7870072ecfa9","VC-Attention低比特注意力:视频生成提速1.59倍","vc-attention-low-bit-video-attention","2026-09-17T13:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"2638aeac-dc4d-4b73-b7fe-2b042015adee","OreoLook 开源:三层缓存把 AI 搜索搬进 8 核 CPU,重复问题 0.1 毫秒出答案","oreolook-three-layer-cpu-cache","2026-09-10T23:08:36+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"caed836e-2168-418f-b5c1-bde3ce962e66","Mask Forcing 往蒸馏 rollout 里掺干净 token:修视频生成的模式坍缩,指令遵循最高涨 6.5 分","mask-forcing-video-diffusion-distillation","2026-09-09T23:08:37+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"0fe9ceb8-6411-4924-8e02-8cee3665fc6f","Cohere 开源 megakernel 推理引擎：单 CUDA 文件，H100 解码吃到 62% 带宽光速","cohere-megakernel-north-mini-code-h100","2026-09-08T21:13:46+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"0c29e1ad-914a-4b79-a153-445c087acb03","被 LLM 抛弃的 dropout 翻身:Cerebras 称调好可省 25% 训练 FLOPs","dont-drop-dropout-layer-sparsity","2026-09-07T21:06:35+00:00"]