[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dreamx-creator-7b-native-audio-video":3,"topics-all":38,"news-related-7ddc323f-fc52-406a-b6df-79b7393e121b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"7ddc323f-fc52-406a-b6df-79b7393e121b","高德开源 DreamX-Creator:7B 原生音视频生成,2K 输出","高德 AMAP-ML 团队发布 DreamX-Creator 1.0:7B 原生联合音视频生成,首帧加文本提示同时产出画面与声音,自回归精炼升至 2K 分辨率,Apache 2.0 许可,模型权重尚未释放。","视频生成模型大多是「哑巴」:画面出来之后,声音要么干脆没有,要么由另一个模型在后面补一道工序。8 月 31 日,高德 AMAP-ML 团队以 DreamX Team 名义在 arXiv 发布 DreamX-Creator 1.0 技术报告,给出另一条路线——用 7B 参数的原生联合音视频生成器,让画面和声音在同一个模型里一起去噪。\n\n## 一天冲上 HF papers 榜首\n\n论文(编号 2608.31106)次日被提交到 Hugging Face Daily Papers,截至发稿已拿到 81 个 upvotes,居当日榜单首位;配套 GitHub 仓库(AMAP-ML\u002FDreamX-Creator)已有 82 stars,采用 Apache 2.0 许可证。研究团队共 10 位作者,第一作者为 Jiashu Zhu,名单中还包括 Xiangxiang Chu。\n\n热度不奇怪。原生音视频联合生成是当前视频生成里少数还没有定型答案的问题:论文摘要直接点出,近来的视频生成器「常常省略音频,或在单独阶段合成音频」,视觉动态与声学事件之间因此缺少相互建模。DreamX-Creator 的回应是把两条流放进同一个网络。\n\n## 架构:前半各走各的,后半门控耦合\n\nDreamX-Creator 以首帧加文本提示为条件。网络前半段,音频流和视频流分别独立处理;后半段通过门控跨模态注意力(Gated Cross-Modal Attention)耦合——每个跨模态注意力头的输出都由 token 级和头级门控调制,官方称由此实现双向音视频交互。\n\n训练侧是四件套:\n\n- **统一音视频数据系统**:构建时序连贯的片段、生成结构化多模态标注、按能力导向组织数据池;\n- **渐进式联合训练**:两个音视频预训练阶段,再加高质量微调;\n- **音视频强化学习**:多模态反馈按视频、音频、跨模态三类路由到对应流,相当于把 RL 后训练从纯文本域搬进音视频域;\n- **自回归 1-Step 2K 精炼**:把一个双向多步教师改造成自回归多步精炼器,再蒸馏成学生模型,每个时间块只需一次去噪评估。\n\n官方口径称,整体性能「可与最强开源系统竞争」。\n\n## 站在 Wan 和 MOVA 的肩膀上\n\nREADME 致谢里点名了两个基础:Wan 团队(Wan2.2)与 OpenMOSS 团队(MOVA)。换句话说,「紧凑 7B 也能做原生音视频」这个故事的前提,是开源视频生成栈已经成熟到可以当基础设施用——后来者不必从底座造起,只需在联合建模、RL 后训练这些增量问题上发力。\n\n## 权重还没放,先读论文\n\n也要泼一盆冷水:GitHub 路线图显示,目前完成的是「初始化仓库」和「发布 1.0 技术报告」两项,第三项——「释放经过验证的模型权重、推理代码、配置和评测工具」——尚未勾选。也就是说,「democratizing」(平民化)当下的真实状态是:许可证已定(Apache 2.0)、配方已公开、权重在路上。想在本地跑起来的开发者,现阶段只能精读技术报告,等下一批交付。\n\n## 所以呢\n\nDreamX-Creator 值得盯住的有两点。其一,它把「音视频一起生成」从旗舰模型的专属能力拉到 7B 量级,论文明确表示将释放紧凑生成器与 2K 精炼器;其二,它的交付节奏正是行业常态的缩影——论文先行、权重后至,开源承诺与可运行工件之间总有时间差。现在能做的,是把报告读完;等权重落地那天,第一时间去验证「官方口径」里的每一条。\n\n参考:arXiv:2608.31106;github.com\u002FAMAP-ML\u002FDreamX-Creator","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.31106","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"567b31c9-1a32-4a0c-9d44-46d1c46fcd10","en","Amap's DreamX-Creator: 7B Joint Audio-Video Generation at 2K","DreamX-Creator 1.0: a 7B native audio-video generator with 2K refinement from Amap's AMAP-ML team. Apache 2.0, tech report out, weights pending.","Most video generators are mute: the visuals come out first, and sound is either missing entirely or bolted on afterwards by a separate model. On August 31, the AMAP-ML team at Amap (publishing as the DreamX Team) posted the DreamX-Creator 1.0 technical report on arXiv, taking the opposite route — a compact 7B-parameter generator that jointly denoises audio and video inside one model.\n\n## Straight to the top of HF Daily Papers\n\nSubmitted to Hugging Face Daily Papers the next day, the paper (arXiv:2608.31106) had collected 81 upvotes as of writing, ranking first on the day's board. The companion GitHub repository (AMAP-ML\u002FDreamX-Creator) sits at 82 stars under an Apache 2.0 license. The team lists ten authors, with Jiashu Zhu as first author and Xiangxiang Chu among the names.\n\nThe attention is not surprising. Native joint audio-video generation remains one of the open problems in video generation: the abstract points out that recent video generators \"often omit audio or synthesize it in a separate stage,\" which limits reciprocal modeling between visual dynamics and acoustic events. DreamX-Creator's answer is to put both streams in one network.\n\n## Architecture: separate early, gated late\n\nConditioned on a first frame and a text prompt, the network processes the audio and video streams independently through the first half, then couples them in the latter half via Gated Cross-Modal Attention — every cross-modal attention head's output is modulated by token-wise and head-wise gates, which the authors say yields bidirectional audio-video interaction.\n\nThe training recipe has four components:\n\n- **A unified Audio-Video Data System** that constructs temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented pools;\n- **Progressive Joint Training**: two audio-video pre-training stages followed by high-quality finetuning;\n- **Audio-Video Reinforcement Learning** with modality-aware multimodal feedback, routing video, audio, and cross-modal signals to the corresponding streams — effectively moving RL post-training from the text-only domain into audio-video;\n- **Autoregressive 1-Step 2K Refinement**: a bidirectional multi-step teacher is adapted into an autoregressive multi-step refiner, then distilled into a student that needs a single denoising evaluation per temporal chunk.\n\nPer the official claim, overall performance is \"competitive with state-of-the-art open-source systems.\"\n\n## Standing on Wan and MOVA\n\nThe README acknowledgements name two foundations: the Wan team (Wan2.2) and the OpenMOSS team (MOVA). In other words, the premise of \"a compact 7B doing native audio-video\" is that the open-source video-generation stack has matured into infrastructure — newcomers don't have to rebuild the base, only push on the incremental problems of joint modeling and RL post-training.\n\n## Weights not out yet — read the paper first\n\nOne cold shower: the GitHub roadmap shows two items checked — repository initialization and the 1.0 technical report — while the third, \"release validated model weights, inference code, configurations, and evaluation tools,\" remains unchecked. So the current honest state of \"democratizing\" is: license settled (Apache 2.0), recipe published, weights on the way. Developers hoping to run it locally can only study the report for now.\n\n## So what\n\nTwo things make DreamX-Creator worth tracking. First, it pulls joint audio-video generation down from flagship-scale exclusivity to 7B, with an explicit commitment to release the compact generator and the 2K refiner. Second, its delivery cadence is the industry norm in miniature — paper first, weights later, with a time gap between the open-source promise and the runnable artifact. The sensible move now: read the report, and when the weights land, verify every line of the official claim yourself.\n\nReferences: arXiv:2608.31106; github.com\u002FAMAP-ML\u002FDreamX-Creator","dreamx-creator-7b-native-audio-video","2026-09-01T13:10:00Z","2026-09-01T13:08:21.634760Z","2026-09-01T13:08:21.634773Z",true,"agent",254,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"9f82c248-0592-421f-9fd8-ebd2100dcaf5","VideoChat3 全开源 4B 视频 MLLM 一次打通四种能力,I3D-ViT 把时空 token 砍掉 16×","videochat3-4b-mllm","2026-07-15T02:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"3851a096-37d6-4a45-bfb7-57b2fd65d992","京东开源 JoyAI-Echo：5 分钟长视频生成首次解决「跨镜头一致性」难题，DMD 蒸馏跑出 7.5× 加速","joyai-echo-jd-5-min-cross-shot-dmd-7-5x","2026-06-12T02:01:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00"]