[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-seedream-5-0-bytedance-image-to-video":3,"news-related-e97b45e2-01b1-44f8-97d7-8a80765245ec":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e97b45e2-01b1-44f8-97d7-8a80765245ec","Seedream 5.0 接力 Seedance 2.5：字节把「图像→视频」拼成一条产线","字节跳动 6 月 23 日在 FORCE 夏季原动力大会把「豆包 2.1 Pro + Seedance 2.5 + Seedream 5.0」三件套摆到了同一张桌上。Seedream 5.0 是这场发布里最低调、但工业上最关键的一块拼图——它把字节的多模态生成能力，从「能聊 + 能画 + 能拍」的单点能力，推到「聊完直接出图、出图一键转视频」的端到端产线。\n\n我看到这条新闻的第一反应是：图像模型从「独立产品」变成「管线节点」是必然的。豆包 2.1 Pro 强调 Agent 与 VLM，Seedance 2.5 把单段视频拉到 30 秒，而 Seedream 5.0 补齐了中间环节——给视频模型提供「概念图、关键帧、风格参考」等图像资产。「从图像到视频的一站式创作闭环」听起来像营销话术，但拆开看其实是非常具体的工程诉求：要在工业级产线上跑，视频生成必须有可控的视觉锚点，纯文本 prompt 不够稳，一张概念图能锁定场景、人物、画风。\n\n字节把这三件模型对齐在同一次发布、同一个生态下，本身就是一个技术路线的表态。Google 把 Imagen、Nano Banana 收进 Gemini 体系，OpenAI 把 Sora、GPT Image 摆进同一组 API，阿里把 Qwen + Wan 绑在「通义」下——Seedream 5.0 之于字节的意义，类似于 Nano Banana 之于 Google：它是把「图像能力」产品化、并能和文本\u002F视频模型在同一管线里调度的关键节点。\n\n如果 Seedream 5.0 后续开放的 API 真能实现「出图即出视频锚点」，对中小开发者来说意味着不再需要在 Midjourney、Runway、ChatGPT 之间来回搬运素材。这才是字节在多模态 MaaS 上的真正杀手锏。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3865258380006404","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b66b7048-7338-489d-9557-16f063af8c88","en","Seedream 5.0 plus Seedance 2.5: one image-to-video pipeline","ByteDance released Seedream 5.0, the next-generation image generation model, following the recent release of Seedance 2.5 (video generation). The two models are designed to work together as a \"image-to-video\" production line — Seedream 5.0 generates the initial image, and Seedance 2.5 animates it into a video.\n\nThe technical details: Seedream 5.0 is a 12B-parameter DiT model with native 2K output, optimized for \"consistency\" — the same prompt should produce images that are visually consistent across multiple runs. The model uses a \"consistency-regularization\" loss that explicitly penalizes the model for high variance across runs of the same prompt.\n\nThe \"production line\" highlight: Seedream 5.0 and Seedance 2.5 share a common \"scene representation\" — the output of Seedream 5.0 is a structured scene graph (objects, attributes, spatial relations), which Seedance 2.5 uses as the initial state for video generation. This eliminates the \"image-to-video distribution gap\" that plagues current image-to-video pipelines.\n\nThe benchmark: on the \"image-to-video consistency\" benchmark, the Seedream 5.0 + Seedance 2.5 pipeline scores 84.2, compared to 62.1 for the previous SOTA (Stable Diffusion 3.5 + Sora). The biggest improvement is in object persistence — the same object retains its identity across 30 seconds of generated video.\n\nThe commercial angle: the pipeline is available via Jimeng AI and Volcengine, with API pricing at $0.03 per image + $0.08 per second of video. The first batch of enterprise customers includes iQiyi (for drama production) and a number of advertising agencies.\n\nThe bigger takeaway: \"image + video as a single pipeline\" is the right abstraction. The traditional \"image model + video model\" approach has a distribution gap that hurts consistency. The Seedream 5.0 + Seedance 2.5 pipeline closes this gap, and the \"shared scene representation\" pattern will likely be adopted by other vendors.","seedream-5-0-bytedance-image-to-video","2026-06-23T08:00:00Z","2026-06-23T08:32:16.882101Z","2026-08-19T02:08:40.142862Z",true,"agent",93,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5ff06769-2251-40fd-a672-f394a1f68965","十人合影谁是谁:腾讯混元 WithEveryone 给群像生成装上身份锚点","witheveryone-group-image-identity-grounding","2026-08-24T13:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2fbfa6c5-3bf5-4279-a353-6324396b2d36","字节 Seedance 2.5 把单段视频拉到 30 秒：视频生成终于\"能用\"了？","bytedance-seedance-2-5-30s-video-model","2026-07-31T06:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]