[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-seedance-2-5-30s-video-model":3,"news-related-2fbfa6c5-3bf5-4279-a353-6324396b2d36":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"2fbfa6c5-3bf5-4279-a353-6324396b2d36","字节 Seedance 2.5 把单段视频拉到 30 秒：视频生成终于\"能用\"了？","字节跳动在火山引擎 FORCE 大会上发布 Seedance 2.5,单段直出 30 秒、50 个多模态参考素材、区域级编辑 + 3D 白膜预览同步上线。30 秒长度、参考素材从 12 提升到 50、4K 原生成片,这三个数字背后是视频生成从\"炫技 demo\"向\"产线可用\"的转变。","# 字节 Seedance 2.5:从 15 秒到 30 秒,视频生成的\"工业可用\"门槛跨过去了\n\n## 30 秒单段直出,不是 PPT 上的数字游戏\n\n字节跳动 7 月 31 日通过 36 氪等渠道正式落地 Seedance 2.5 的企业接入,这款豆包系视频生成模型在 6 月 23 日火山引擎 FORCE 大会上首次发布时,最扎眼的数字是 **30 秒**——单次生成的连续视频时长直接翻倍到上一代(Seedance 2.0)的 2 倍,且不是\"几段拼起来\",而是一条原生直出。\n\n这事的难度被很多评测低估了。视频模型在 5-8 秒内还能维持人物一致性、动作合理性和镜头逻辑,但只要拉长,人物脸会漂、光照会跳、动作会失真。市面上绝大多数\"30 秒视频\"其实是把 4-6 段 5 秒短片拼起来,接缝处肉眼可见。Seedance 2.5 这次如果真做到 30 秒不拼接,意味着模型内部的时序一致性建模有了实质性突破——而不仅仅是把推理窗口拉长。\n\n## 三大升级:长度、参考、可控性\n\n发布会上官方把升级点压成三件事:\n\n1. **30 秒单段直出**(2.0 是 15 秒)——生产意义上的\"可叙事单元\",漫剧、广告分镜、TVC pre-vis 都能用\n2. **单项目 50 个多模态参考素材**(2.0 是 12 个)——人物设定、产品图、风格帧、品牌色都能塞进一次 prompt\n3. **区域级编辑 + 3D 白膜预览**——前者在不动原片运镜、灯光的前提下替换主体;后者把分镜表从\"靠猜\"变成\"先排,后生成\"\n\n另外原生 4K 也跟上了,但要小心:官方主视觉把 4K 写在 Seedance 2.0\u002F2.5 共用能力里,**不是 2.5 独占**。Tosea 这条比较容易混淆,大家看的时候要分清楚。\n\n## 为什么是现在:2.0 已经坐在 AI 视频榜首\n\n任何\"2.X\"升级都不能脱离上一代的表现谈。在 Artificial Analysis 的 Text-to-Video Arena(盲测人类偏好)上,Seedance 2.0(Dreamina 标签下)以 **Elo 1219 排第一**,领先 Kling 3.0 Pro(1105)、Google Veo 3.1(1094)。这是 7 月 1 日的公开数据,不是发布会 PPT。\n\n换句话说,字节这次发 2.5 是在**已经在榜首**的位置上做的增量,不是追赶。而且参照 2.0 的标准化定价(每分钟 1080p 视频约 $9,而 Veo 3.1 约 $24、Kling 3.0 Pro 约 $20),2.5 上市后的核心卖点大概率还是\"价格一半,质量不输\"——这是 Veo\u002FKling 这类海外模型最难跟的维度。\n\n## 三个被低估的细节\n\n**(1)50 个参考素材的真正意义**——不是\"素材库更大\",而是把\"角色一致性\"和\"品牌一致性\"从**事后修**变成**事前锁**。广告场景里,一个产品的 logo、模特、包装、配色能否在 30 秒里保持一致,是 AI 视频能不能进产线的生死线。50 个参考就是给模型 50 个\"约束条件\"。\n\n**(2)区域级编辑的产业意义**——这条对中小广告公司最有用。原来换市场需要重新生成整条片,现在只换主体,运镜和灯光保持不变。对一个出海品牌做 4 个区域版本,工作量从\"4 条新片\"变成\"1 条片 + 3 次局部编辑\"。\n\n**(3)3D 白膜预览**——这是给\"导演\"用的工具,不是给\"剪辑师\"。在生成前先用粗 3D 模型排镜头走位,生成时按这个走位去合成。**这意味着 AI 视频开始侵入 pre-vis 流程**,而 pre-vis 一向是好莱坞和 4A 广告公司的\"前置成本黑洞\"。\n\n## 我的判断:这波不是噱头,但要冷静看\n\n值得兴奋的:\n- 30 秒单段直出如果实测成立,意味着 AI 视频跨过\"单镜头可用\"门槛,进入\"多镜头叙事\"阶段\n- 50 个参考 + 区域级编辑,把\"一致性\"和\"可迭代\"两个产线核心痛点都解了\n- 价格战层面,字节对 Veo\u002FKling 仍然有显著成本优势\n\n要冷静的:\n- 7 月初才正式上线,**目前所有 2.5 性能数据都是官方自报**,独立盲测要等至少 2 周\n- 50 个参考在\"角色一致性\"上的实际效果,要等真实广告团队跑出来\n- 4K 原生在 2.0 也能用,不要被\"4K 升级\"的话术带偏\n\n## 对国内视频生成赛道的冲击\n\n可灵(Kling)、Vidu、PixVerse 这些玩家要在 2-3 个月内回应。字节这次发布的**不是单点参数升级**,而是把\"长度 + 一致性 + 可控性 + 价格\"四件套一起打出来——任何只回应其中一项的友商,都会显得被动。\n\n更长远看,**Seedance 2.5 + Seedream 5.0 + Seed-Audio 1.0 + Doubao 2.1 Pro** 同步在 FORCE 上发出来,字节其实在拼\"全模态产线\"——这是 2024 年 OpenAI 用 Sora\u002FDALL-E\u002FWhisper 想做但没做完的事。国内厂商现在比海外同行更接近\"一站式生成套件\"的产品形态。\n\n## 所以呢?\n\n如果你是做广告\u002F短剧\u002F漫剧的内容团队,**7 月中旬之后值得花一周时间做一次内部 PoC**——拿你最头疼的\"多市场版本\"或\"长镜头一致性\"两个场景,跑一下 Seedance 2.5 对比当前 SOTA,基本能判断这条产线值不值得切。\n\n如果你是投资人,关注的不是 Seedance 2.5 本身,而是**字节视频生成定价策略对整个赛道毛利率的压力**——参考 2.0 的 $9\u002Fmin,2.5 大概率维持,甚至更低。Kling\u002FVidu 的海外客单价会被进一步压缩。\n\n如果你是技术人,关注 30 秒直出的训练方法论——这是公开数据里几乎没人拆过的点,大概率字节在 **时序一致性 loss + 滑动窗口 attention + 数据混合比例** 三个地方下了真功夫。等技术报告出来再细看。\n\n(本文基于火山引擎 FORCE 大会公开信息、Seedance 2.0 在 Artificial Analysis 的公开 Elo 数据,以及快科技、tosea.ai 等媒体报道整理)","https:\u002F\u002Fnews.mydrivers.com\u002F1\u002F1131\u002F1131453.htm","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3a69b722-6ede-49e3-9de0-4dadb9d8d454","en","Seedance 2.5 stretches single clips to 30 seconds","ByteDance launched Seedance 2.5 at the Volcano Engine FORCE conference: 30-second single-clip output, up to 50 multimodal reference materials, region-level editing and 3D white-model previsualization. Doubling the single-take length from 15 to 30 seconds, raising the reference cap from 12 to 50, and adding native 4K together mark the shift of AI video from \"showy demo\" to \"production usable.\"","# ByteDance Seedance 2.5: Pushing Single-Shot Video to 30 Seconds — Is AI Video Finally \"Production-Ready\"?\n\n## 30 seconds, native single-take, and why that number is not marketing fluff\n\nByteDance's Doubao-family video generation model Seedance 2.5 was first unveiled on June 23, 2026 at the Volcano Engine FORCE conference in Beijing, and as of late July 2026 it is now rolling out to enterprise customers via platforms like Jimeng AI, Doubao Pro and Volcano Ark. The headline number is **30 seconds** — the maximum length of a single, continuously generated video clip has been doubled from Seedance 2.0's 15-second ceiling, and the model produces it as one native generation, not as a stitched-together sequence of shorter segments.\n\nThe difficulty here is consistently underestimated. Video models can hold character consistency, physical motion and camera logic for 5–8 seconds. Push past that and faces drift, lighting shifts, motion loses plausibility. Most \"30-second videos\" in the market today are actually 4–6 stitched 5-second clips with visible seams. If Seedance 2.5 genuinely produces 30 seconds without stitching, the model has made a real architectural breakthrough in long-horizon temporal consistency — not just a longer inference window.\n\n## The three pillars: length, references, controllability\n\nThe official launch compresses the upgrade into three claims:\n\n1. **30-second single-take output** (vs 15 seconds on 2.0) — a \"narrative unit\" that can be used in short-form drama, advertising storyboards and TVC pre-visualization\n2. **Up to 50 multimodal reference materials per project** (vs 12 on 2.0) — character designs, product shots, style frames, brand colors all go into a single prompt\n3. **Region-level editing + 3D white-model previsualization** — the former replaces subjects without disturbing the original camera move, lighting or motion; the latter lets creators block out shots in a rough 3D scene before committing to a full generation\n\nNative 4K output is also supported, though with an important caveat: ByteDance's own headline slide puts 4K as a **shared** 2.0\u002F2.5 capability, not a 2.5 exclusive. Some third-party coverage has muddled this — read carefully.\n\n## Why now: because 2.0 is already sitting on top\n\nAny \"2.X\" launch should be evaluated against the prior generation's standing. On Artificial Analysis's Text-to-Video Arena (blind human preference), the shipping Seedance 2.0 (under the \"Dreamina\" label) sits at **Elo 1219 — first place**, ahead of Kling 3.0 Pro (1105) and Google Veo 3.1 (1094). These are public numbers from the start of July, not the launch keynote.\n\nIn other words, ByteDance is shipping 2.5 from a position of strength, not catching up. And given 2.0's normalized pricing (~**$9 per minute** of 1080p video, vs ~$24 for Veo 3.1 and ~$20 for Kling 3.0 Pro), the 2.5 pitch is likely to remain the same: same-or-better quality at roughly half the price. That is the hardest dimension for overseas competitors to match.\n\n## Three under-appreciated details\n\n**(1) What 50 reference materials actually means** — not \"more assets in the library\" but moving \"character consistency\" and \"brand consistency\" from **post-hoc fixing** to **prior-time locking**. In advertising, whether a product's logo, model face, packaging and color palette stay locked across 30 seconds of footage is the line between \"AI video\" and \"usable AI video.\" Fifty references give the model fifty constraints to honor.\n\n**(2) The industrial significance of region-level editing** — this matters most for small-to-mid advertising teams. Localizing a campaign used to mean re-generating the entire spot. Now you swap the subject, the on-screen talent, the packaging copy, and the camera move, lighting and motion are preserved. For a brand running four regional versions, the workload shifts from \"four new spots\" to \"one spot + three local edits.\"\n\n**(3) 3D white-model previsualization** — a tool built for the **director**, not the editor. Block out camera moves in rough 3D before generation; the model then synthesizes along that blocking rather than letting the camera language emerge by trial and error. This means AI video is starting to invade the **pre-vis** workflow — historically one of the most expensive front-loaded cost sinks in Hollywood and 4A agency production.\n\n## My read: this is real, but stay cool\n\nWorth being excited about:\n- If 30-second single-take holds up in independent testing, AI video crosses the \"single-shot usable\" threshold and enters \"multi-shot narrative\" territory\n- 50 references + region-level editing address the two core production pain points (consistency and iteration)\n- On the price dimension, ByteDance still has a structural cost advantage over Veo and Kling\n\nReasons to stay measured:\n- The 2.5 public release is only just happening in early July; **every performance number so far is vendor-reported** — independent blind tests will take at least 2 weeks\n- The real-world effectiveness of 50 references on character consistency needs to be tested with actual advertising briefs\n- 4K native is also on 2.0; do not be led by \"4K upgrade\" marketing language\n\n## The impact on China's video generation race\n\nKling (Kuaishou), Vidu, PixVerse and other domestic players now have a 2–3 month window to respond. ByteDance is not playing a single-axis parameter game here — they are hitting **length + consistency + controllability + price** as a package. Any competitor that responds on only one of those four axes will look reactive.\n\nLonger term: Seedance 2.5, Seedream 5.0, Seed-Audio 1.0 and Doubao 2.1 Pro were all announced together at FORCE. ByteDance is assembling a **full-modality production line** — the same ambition OpenAI tried to execute in 2024 with Sora \u002F DALL-E \u002F Whisper but never fully shipped. Chinese vendors are now closer to a \"one-stop generative suite\" product shape than their Western peers.\n\n## So what?\n\nIf you run an ad \u002F short-drama \u002F 漫剧 content team, **a one-week internal PoC after mid-July is worth the time** — pick the two scenarios you currently find most painful (multi-market versioning, long-take consistency) and benchmark Seedance 2.5 against your current SOTA. That is enough signal to decide whether to switch the production line.\n\nIf you are an investor, the real story is not Seedance 2.5 itself but **the pricing pressure ByteDance is putting on the entire category's gross margin** — referencing 2.0's $9\u002Fmin, 2.5 will likely match or undercut. Kling and Vidu's overseas unit pricing will be further compressed.\n\nIf you are a technical practitioner, watch for the training methodology behind 30-second single-take output — this is a point that almost nobody has publicly dissected. ByteDance almost certainly did real work on **temporal consistency loss + sliding-window attention + data mixture ratios**. Wait for the technical report.\n\n(Based on public information from the Volcano Engine FORCE conference, Seedance 2.0's standing on the Artificial Analysis leaderboard, and reporting from MyDrivers and tosea.ai)","bytedance-seedance-2-5-30s-video-model","2026-07-31T06:00:00Z","2026-07-31T08:05:07.288426Z","2026-07-31T08:05:07.288434Z",true,"agent",105,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"f0ea091f-d030-4a32-827a-ae21c1b61c8f","昆仑万维 WAIC 大会将一次性放出四款全模态模型：Matrix-3.5 把\"状态-动作\"塞进一套参数","kunlun-waic-matrix-3-5","2026-07-17T12:08:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00"]