[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-seedance-2-unit-dit":3,"news-related-ba4fec9d-1a6e-49db-9669-1e4b168afca2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ba4fec9d-1a6e-49db-9669-1e4b168afca2","字节 Seedance 2.0 翻身仗：一次从 UNet 到 DiT 的架构选择","字节在大模型上的\"第一场翻身仗\"，36 氪把它归到了 Seedance 2.0 头上。背后是一个并不复杂、但少有人讲清楚的技术判断——把视频生成的训练架构，从 2D UNet 扩展到 3D，改成以 DiT（Diffusion Transformer）为基底的原生视频路线。\n\n早期字节内部 AI Lab 的 PixelDance 走的是\"图像扩散视频化\"捷径：用 2D UNet 扩成 3D，结构上更快、更稳，但上限被图像模型绑死。同一时间，可灵和 Sora 都已转向 DiT，字节因此在视频生成上落后近一年。\n\n从 PixelDance 后期过渡到 Seedance 时，最大的变化就是把架构换到了 DiT 基底。DiT 的好处是更贴 Scaling Law——参数量、数据量、算力继续变大，效果能持续涨。Seedance 2.0 把参数量推到 200B 级别，被多位业内人士评价为\"模型足够大 + 数据足够丰富\"两条线的合流。\n\n更隐形的是数据。Seedance 每个算法背后配着十数位数据同事，背后是一个上千人的评测团队，靠细致标注把用户提示词精准匹配到训练数据上。素材大多采买自影视级镜头，再让 LLM 拆解成脚本和分镜。字节内部对所有模型都有\"不做蒸馏\"的共识——目标定为全球 SOTA，必须自己合成、清洗数据。\n\n商业上的转折也更陡。火山引擎今年 MaaS 收入中 Seedance 已经贡献过半，720P 视频定价 1 元\u002F秒仍不打折。Seedance 2.5 将于 7 月下旬发布，火山目标非常清晰——在 Veo 之后挤进全球第一。\n\n所以当字节把豆包 2.1 拉上 Coding\u002FAgent 牌桌时，视频模型这边已经跑通了\"模型好→场景好→利润高\"的闭环。这门被低估的生意，才刚刚开始。","https:\u002F\u002F36kr.com\u002Fp\u002F3885177884078083","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ddb3c2bd-3b6e-4bf3-8cde-5187bf731ce1","en","Seedance 2.0's comeback: the UNet-to-DiT architecture bet","ByteDance's \"first comeback\" in the large-model arena, 36Kr attributes it to Seedance 2.0. Behind it is a not-very-complex but rarely clearly explained technical judgment — moving the training architecture of video generation from 2D UNet extension to 3D, into a native video pipeline with DiT (Diffusion Transformer) as the base. Early ByteDance's AI Lab's PixelDance took the \"image diffusion to video\" shortcut: using 2D UNet extended to 3D, structurally faster and more stable, but with the upper bound locked by the image model. At the same time, Kling and Sora had already shifted to DiT, leaving ByteDance about a year behind on video generation. When transitioning from late PixelDance to Seedance, the biggest change was switching the architecture to a DiT base. DiT's benefit is closer to the Scaling Law — as parameters, data volume, and compute continue to grow, the effect continues to climb. Seedance 2.0 pushes parameters to the 200B tier, evaluated by multiple industry insiders as the confluence of the \"model is large enough + data is rich enough\" two lines. More invisible is the data. Behind every Seedance algorithm sit over a dozen data colleagues, and behind them a thousand-person evaluation team, who use meticulous annotation to accurately match user prompts to training data. Most material is bought from film-grade footage, then decomposed by LLM into scripts and storyboards. ByteDance internally has a \"no distillation\" consensus on all models — the goal is global SOTA, and they must synthesize and clean their own data. The commercial turning point is even sharper. Among Volcano Engine's MaaS revenue this year, Seedance has already contributed more than half, with 720P video priced at 1 yuan\u002Fsecond with no discount. Seedance 2.5 will be released in late July, and Volcano's goal is very clear — squeeze into global first place after Veo. So when ByteDance pulls Doubao 2.1 onto the Coding\u002FAgent table, the video model side has already run through the closed loop of \"good model → good scenarios → high profit\". This underestimated business has only just begun.","bytedance-seedance-2-unit-dit","2026-07-08T00:30:00Z","2026-07-08T00:03:58.414876Z","2026-08-19T02:08:40.142862Z",true,"agent",119,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5ff06769-2251-40fd-a672-f394a1f68965","十人合影谁是谁:腾讯混元 WithEveryone 给群像生成装上身份锚点","witheveryone-group-image-identity-grounding","2026-08-24T13:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00"]