[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-puffin-world-native-3d-world-states":3,"topics-all":42,"news-related-dc2f4ead-963c-4a8e-bd41-400bebf83bb4":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":28,"news_slug":35,"published_at":36,"created_at":37,"modified_at":38,"is_published":39,"publish_type":40,"image_url":15,"view_count":41},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","NTU S-Lab 等机构发布 Puffin-World:统一多模态架构同时建模物理、几何、外观三种原生世界状态,配套 Puffin-16M 数据集(15M 三元组+1M 轨迹)。四个相机感知基准全部拿下中位误差第一,代码、模型、数据集已开源。","世界模型大多只在像素层面打转:画面越来越精致,可一旦相机做横滚、俯仰这类非常规运动,重力方向漂了、几何塌了,穿帮立现。S-Lab 南洋理工、密歇根大学、北京交通大学与 ACE Robotics 的联合团队换了条路:与其事后修补,不如把物理和几何当作模型的原生状态([arXiv:2609.04196](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.04196))。\n\n## 三种原生世界状态\n\nPuffin-World 是一个统一多模态架构,同时建模三类互补的世界状态:物理(重力场与纬度)、几何(深度)、外观(图像)。架构上由几何对齐的视觉编码器、LLM、扩散模型和轻量连接器组成,理解和生成共用同一套参数——同一个模型既能解读相机几何,也能合成世界一致的观测,不依赖任务专属的外部几何模块。\n\n两处设计撑起了这条路线:一是 Omni-Camera 表征,用重力感知的绝对透视场配射线式相对几何,9 通道条件精确控制相机内参、朝向与轨迹;二是物理传播策略,把物理约束沿生成轨迹向前传递,在极端旋转和长相机运动下保住重力一致性。\n\n## 四个基准的中位误差全第一\n\n项目页报告称,在 Stanford2D3D、MegaDepth、TartanAir、LaMAR 四个公开基准上,Puffin-World 拿下 12\u002F12 项最佳中位误差与 33\u002F36 项最佳 AUC;LaMAR 上 roll 误差低至 0.26°,Stanford2D3D 上 vFoV 误差 1.62°。生成侧,自建 Puffin-Cam-Bench 上 up-vector \u002F latitude \u002F gravity 中位误差分别为 0.84°\u002F1.26°\u002F0.79°,FID 最低;RealEstate10K 上 PSNR 17.22、LPIPS 0.318 双指标第一;Puffin-Traj-Bench 上 roll\u002Fpitch 中位误差 0.80°\u002F1.10°。闭环应用给了两类演示:模仿式世界探索与自校准探索——后者由模型推理当前物理状态并预测修正相机动作。\n\n## 44.5M 张相机标注图怎么来\n\n Scaling 是关键:Puffin-16M 数据集由 1500 万视觉-语言-相机三元组和 100 万条多样相机轨迹组成,团队还为 28 个常用公共数据集标注了 roll\u002Fpitch\u002FvFoV,合计约 4450 万张相机标注图像,全部随 Hugging Face collection 放出。代码、模型、数据集三层全开源。\n\n## 该泼的冷水\n\n对照表里藏着细节:vFoV 这个维度上,专用相机内参估计器 AnyCalib 在 LaMAR 的中位误差 2.25°,仍优于 Puffin-World 的 2.73°——专材模型在特定指标上没有退场。另外 12\u002F12 这类「全第一」是团队在自家项目页上对选定基线组合的报告,接收侧尚未有独立复现;统一架构做多任务,每个单点能力是否都扛得住专模型的挤压,要等第三方评测说话。\n\n但对做世界模型的人来说,方向信号是明确的:像素保真之外,把重力、深度这些物理量当作一等公民建模,再配上 4450 万张标注图的开源底座,空间智能这条线从「会画」往「懂物理」又推了一步。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.04196","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,16,19,22,25],{"id":12,"name":13,"slug":13,"description":14,"color":15},"9112951a-2abb-4214-b63a-385ec7afb2ba","ai-for-science","AI for Science 专题：追踪 AI 在生命科学、化学材料、物理世界模型等科学方向的关键突破",null,{"id":17,"name":18,"slug":18,"description":15,"color":15},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":20,"name":21,"slug":21,"description":15,"color":15},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":23,"name":24,"slug":24,"description":15,"color":15},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":26,"name":27,"slug":27,"description":15,"color":15},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[29],{"id":30,"lang":31,"title":32,"summary":33,"content":34},"aa85a8c8-2d8d-40ac-ae0c-31db5460bca0","en","Puffin-World: One Open Model for Physics, Geometry, Appearance","One multimodal model jointly learns physics, geometry, and appearance world states; ships Puffin-16M data and takes 12\u002F12 best median errors.","Most world models stay at the pixel level: frames keep getting prettier, but once the camera performs unconventional motions like roll or pitch, gravity drifts, geometry collapses, and the illusion breaks. A joint team from S-Lab at Nanyang Technological University, the University of Michigan, Beijing Jiaotong University, and ACE Robotics took a different path: instead of patching afterwards, treat physics and geometry as native states of the model ([arXiv:2609.04196](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.04196)).\n\n## Three Native World States\n\nPuffin-World is a unified multimodal architecture that jointly models three complementary world states: physics (gravity field and latitude), geometry (depth), and appearance (image). Architecturally it combines a geometry-aligned vision encoder, an LLM, a diffusion model, and a lightweight connector, with understanding and generation sharing the same parameters — the same model can interpret camera geometry and synthesize world-consistent observations, without task-specific external geometry modules.\n\nTwo designs hold the line: the Omni-Camera representation pairs a gravity-aware absolute perspective field with ray-based relative geometry, using a 9-channel condition for precise control over camera intrinsics, orientation, and trajectory; and a physics propagation strategy carries physical constraints forward along generated trajectories, preserving gravity consistency under extreme rotations and long camera motion.\n\n## First on Median Error Across Four Benchmarks\n\nThe project page reports that on Stanford2D3D, MegaDepth, TartanAir, and LaMAR, Puffin-World takes 12\u002F12 best median-error results and 33\u002F36 best AUC metrics; roll error reaches 0.26° on LaMAR and vFoV error 1.62° on Stanford2D3D. On the generation side, its self-built Puffin-Cam-Bench shows median up-vector \u002F latitude \u002F gravity errors of 0.84° \u002F 1.26° \u002F 0.79° with the lowest FID; on RealEstate10K it ranks first in both PSNR at 17.22 and LPIPS at 0.318; and on Puffin-Traj-Bench its median roll\u002Fpitch errors are 0.80° \u002F 1.10°. For closed-loop applications the team demonstrates mimic world exploration and self-calibrated exploration, where the model reasons about the current physical state and predicts corrective camera actions.\n\n## Where 44.5M Camera-Labeled Images Come From\n\nScaling is the key: the Puffin-16M dataset comprises 15 million vision-language-camera triplets and 1 million diverse camera trajectories. The team also annotated roll\u002Fpitch\u002FvFoV for 28 widely used public datasets — roughly 44.5 million camera-grounded images in total — all released through a Hugging Face collection. Code, models, and datasets are all open-sourced across three layers.\n\n## The Cold Water\n\nDetails hide in the comparison table: on vFoV, AnyCalib, a specialized camera-intrinsic estimator, still beats Puffin-World on LaMAR with a 2.25° median error versus 2.73° — specialized models have not exited the stage on specific metrics. And \"12\u002F12 first-place\" figures come from the team's own project page against their chosen baseline set, with no independent replication yet; whether every single capability of a unified multi-task architecture withstands pressure from specialized models will take third-party evaluation.\n\nFor world-model builders, though, the directional signal is clear: beyond pixel fidelity, modeling gravity and depth as first-class citizens — plus an open-sourced base of 44.5 million annotated images — moves spatial intelligence another step from \"can paint\" toward \"understands physics\".","puffin-world-native-3d-world-states","2026-09-06T19:09:41Z","2026-09-06T19:09:49.377748Z","2026-09-06T19:09:49.377759Z",true,"agent",159,[43,51],{"slug":13,"tag_slug":13,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":39,"created_at":49,"modified_at":50},"AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":39,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"630ed9ae-699e-4115-8e75-33196ea6db28","MiniMax Music 3 开源:8B+0.6B 双 LLM 写五分钟完整歌,8GB 显存能跑","minimax-music3-open-weights-architecture","2026-08-29T13:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b1400260-ba9f-4e84-b658-ce53abba9304","BreezeBlue 开源 Breeze TTS 2:3B 参数实时语音,五语种、可控设计、首包 133 毫秒","breeze-tts-2-open-source-realtime","2026-08-29T10:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"d8bc7b5e-9eb0-475e-91b7-5a3390d2c6a6","2026年开源LLM爆发：Meta、阿里、Google竞相发布新一代模型","open-source-llm-boom-2026-q1-meta-alibaba-google","2026-04-24T04:06:08+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"108af093-b226-4372-9cf0-77323ffc5456","小鹏 X-AuT 给语音大模型剪枝:音频塔砍 4 层,车载推理提速 21.4%","xpeng-x-aut-audio-encoder-pruning","2026-09-12T19:06:47+00:00"]