[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-bernini-mllm-semantic-planner-diT":3,"news-related-ff7fa3e1-6737-4bd9-85c8-8010d13a44f3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ff7fa3e1-6737-4bd9-85c8-8010d13a44f3","字节跳动 Bernini 开源：用 MLLM 当\"语义规划师\"，拆开视频生成的\"思考\"与\"渲染\"","字节跳动 Bernini 团队在 arXiv 发布《Bernini: Latent Semantic Planning for Video Diffusion》，提出把 MLLM 与扩散模型在视频生成中显式分工的统一框架：MLLM 负责\"语义规划\"，扩散模型负责\"像素渲染\"。\n\nBernini 把\"语义表示\"显式定义在 ViT 嵌入空间，规划器输出的语义向量可被 DiT 渲染器直接作为条件输入，规避文本瓶颈。两模块可独立训练再轻量协同，兼顾 MLLM 理解力与 DiT 像素质量。配合 Segment-Aware 3D RoPE 与规划器内 chain-of-thought，Bernini 在多个视频生成与编辑 benchmark 取得 SOTA，Hugging Face 已开源 Bernini-R（Apache 2.0）。\n\n这是为下一代视频生成系统定义\"操作系统级\"接口——MLLM 决定\"做什么、为何做\"，DiT 决定\"如何画\"。Sora、可灵、Wan 把参数堆到百亿量级时，行业真正欠缺的或许不是更大的渲染器，而是一条更清晰的\"语义 ↔ 像素\"对接通道。Bernini 正在填补它。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.22344","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ba0827a7-cc43-4316-94a1-4b1eacd3d51f","en","ByteDance's Bernini splits video thinking from rendering","arXiv 2605.22344 introduces Bernini, an open-source video generation framework from ByteDance that uses an MLLM (Multimodal Large Language Model) as a \"semantic planner\" to separate the \"thinking\" from the \"rendering\" in video generation. The standout: the semantic planner decides \"what should happen in the video,\" and a separate rendering model generates the actual pixels.\n\nThe \"semantic planner + renderer\" insight: traditional video generation models do both \"thinking\" and \"rendering\" in a single model. The \"thinking\" — deciding what objects to show, what actions to take, what the narrative is — is conceptually different from \"rendering\" — converting the decision into pixels. Bernini's fix: separate the two, with a \"semantic planner\" (an MLLM) handling the \"thinking\" and a \"renderer\" (a video diffusion model) handling the \"rendering.\"\n\nThe technical details: the semantic planner takes a text prompt and outputs a structured \"scene description\" — a sequence of (object, action, time) triples that describe the video. The renderer takes the scene description and generates the video. The two are trained jointly, with the semantic planner's output being a \"soft constraint\" on the renderer's generation.\n\nThe benchmark: Bernini-generated videos score higher on \"narrative consistency\" and \"object persistence\" than single-model video generation. The \"semantic planner\" gives the model a higher-level understanding of the video, and the \"renderer\" focuses on visual quality.\n\nThe bigger takeaway: \"semantic planning\" is a significant new direction for video generation. The \"one model does everything\" approach has hit a quality ceiling, and the \"semantic planner + renderer\" approach is a clean way to break through. For the industry, this signals that the next round of video generation models will adopt the \"planner + renderer\" architecture, and the \"best video model\" will be the one with the best semantic planner.","bytedance-bernini-mllm-semantic-planner-diT","2026-05-21T00:00:00Z","2026-06-12T14:38:55.915898Z","2026-08-19T02:08:40.142862Z",true,"agent",139,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b7e375b3-6949-4ff6-9cd4-012aced6bbc3","Runway赌注视频生成通往世界模型：挑战Google的架构之赌","runway-video-world-model-google-bet","2026-05-16T04:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"faab4a6c-9cb0-4f2a-a5bf-1f122306008b","Wan-Streamer v0.2：分辨率 192p→640p，保住 200ms","alibaba-wan-streamer-v0-2","2026-07-10T16:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"aeb60d9a-6639-4669-96a4-951aadad40cb","AI 视频工具进入「全场景」分化期:6 款主流产品的技术路线对比","ai-video-tools-comparison","2026-07-08T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a2e8ac5b-ca51-4ddb-88d4-54373d1f0774","SUNTA 用\"惊奇度\"切分视频预测:东京大学让模型在 250 步后仍不崩溃","sunta-surprise-chunking-video","2026-07-04T16:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2657cbe0-7743-43f2-9332-ee18b84b1229","Directing the World: 中国电信 TeleAI 把自回归视频世界模型推到\"组合控制\"","teleai-directing-the-world","2026-07-01T10:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"efa5f558-e27f-4889-bf71-74cbef874ace","即梦 AI 把 Seedance 2.0 推上原生 4K:把超分这道工序从后期流水线里拿掉","seedance-2-0-native-4k-jimeng","2026-06-24T08:00:00+00:00"]