[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-graphvid-multi-subject-video":3,"news-related-12a49e02-0374-40bd-8ce6-35695e3f19e2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"12a49e02-0374-40bd-8ce6-35695e3f19e2","GraphVid把视频控制从Prompt改成交互图：多主体生成终于有了结构化接口","一段文字让视频里的五个角色同时打架、递东西、绕开障碍，今天的生成模型往往还是会\"各动各的\"。GraphVid 给出的答案，不是继续堆 prompt，而是把场景先画成一张交互图。论文提出 graph-conditioned image-to-video 模型：把对象、动作和对象之间的关系结构化，再交给生成模型执行；同时配套构建 GraphVid-Bench，用带关系标注的视频训练和评估多主体交互。结果很直接：相比 Motion-I2V，FID 最多下降 39.9%，FVD 下降 37.6%，PSNR 从 9.87 提升到 15.98，SSIM 从 0.38 提升到 0.61，而且训练数据和可训练参数都更少。它的价值不只是画面更清晰，而是让\"谁影响谁\"成为模型可计算的条件。遮挡、重叠、多主体协同这些最难靠文字说清的场景，图结构比一长串 prompt 更像导演的分镜表。更重要的是，这种接口把视频生成从\"描述画面\"推向\"编排关系\"：未来的广告、游戏过场和机器人训练数据，都可能先编辑一张场景图，再让模型补出连续画面。当然，论文是预印本，指标依赖数据集和实验设置，离通用生产工具还有距离。但方向已经很明确：视频生成下一阶段拼的不是会不会动，而是能不能按结构动。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.21580","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"8defbba8-a637-485a-937b-2cb7692bb5d1","en","GraphVid swaps prompts for interaction graphs in video control","When a single line of text asks the five characters in a video to fight, hand things over, and dodge obstacles at the same time, today's generation models still tend to \"each do their own thing\". GraphVid's answer isn't more prompts — it's drawing the scene as an interaction graph first. The paper proposes a graph-conditioned image-to-video model that structures the objects, the actions, and the relationships between them, then hands the structured representation to the generator. A companion GraphVid-Bench is built, with relationship-annotated videos used to train and evaluate multi-subject interaction. The results are direct: versus Motion-I2V, FID drops by up to 39.9%, FVD drops by 37.6%, PSNR climbs from 9.87 to 15.98, and SSIM climbs from 0.38 to 0.61 — all with less training data and fewer trainable parameters. The value isn't just sharper frames; it's making \"who influences whom\" a computable condition for the model. In occlusion, overlap, and multi-subject coordination — the cases hardest to describe in text — a graph structure reads more like a director's storyboard than a long prompt. More importantly, this interface pushes video generation from \"describing a picture\" to \"orchestrating relationships\": future ads, game cutscenes, and robot training data may all start by editing a scene graph, then letting the model fill in the continuous frames. The paper is still a preprint, the metrics depend on the dataset and experimental setup, and it's still far from a general-purpose production tool. But the direction is clear: the next stage of video generation isn't about whether it can move — it's about whether it can move according to a structure.","graphvid-multi-subject-video","2026-07-27T00:00:00Z","2026-07-26T22:23:21.051532Z","2026-08-19T02:08:40.142862Z",true,"agent",82,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f47e8ce0-08eb-4e76-966e-7fa47ea64440","DiffusionBench：21 个扩散 Transformer 告诉你，ImageNet 跑分表已经骗了行业好几年","diffusionbench-nanogen-21-dit-imagenet-fid","2026-06-25T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"8f4e0907-919f-410c-b543-7b52260659c2","「Thinking with Video」把推理拉出文本：Sora-2 在 MATH 跑到 92%，多模态统一架构有了新候选","thinking-with-video-sora-2-fudan-92-math","2026-06-16T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f9a2ceff-1bea-4f64-b327-363ca7b2d767","LLM何时会将关键信息「视而不见」？三年研究揭示上下文位置的惊人影响力","llm-context-position-middle-bias-50-models","2026-05-28T14:06:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"95b04c15-d5e5-4dca-ab2b-14e343bdd4e6","UC Berkeley 曝光 AI 基准测试系统性漏洞：45 种方法可在 13 个主流榜单上「不解决任何问题拿满分」","uc-berkeley-benchmark-45-cheats-13-leaderboards","2026-05-15T01:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00"]