[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tempcloze-video-llm-temporal-alignment":3,"topics-all":35,"news-related-ee7c1b35-e8cc-41e3-8b8f-27a512a9f639":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"ee7c1b35-e8cc-41e3-8b8f-27a512a9f639","TempCloze 视频「完形填空」:31 款 Video-LLM 横评,开源模型时间对齐平均 26.54% 逼近乱猜","港大团队推出视频完形填空基准 TempCloze:给模型看开头和结尾,从四个候选里挑出真实中段。31 款 Video-LLM 横评显示时间对齐是最大短板,开源模型平均 26.54%、贴近 25% 随机线,人类基线 97%。","视频模型这两年在「看懂画面」上进步飞快,但港大团队 9 月放出的 TempCloze 基准给整个行业泼了盆冷水:把一段视频掐头去尾,让模型从四个候选片段里挑出真正缺失的中段——这个人类随手就能完成的\"视频完形填空\",把 31 款主流 Video-LLM 考了个底朝天([arXiv:2609.01515](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.01515))。\n\n## 考题怎么出:同源干扰项设计\n\nTempCloze 收录 1,521 条视频,来自 LVD-2M、EgoLife 等 7 个公开来源,以长镜头和第一人称视角为主。每条视频被切成开头(B)、缺失中段(M)、结尾(E)三部分,模型拿到 B 和 E,要从 4 个候选里认出真正的 M。\n\n巧思全在干扰项上,而且同一维度用同源视频构造:语义维度(Semantic)的干扰是同一条视频里互不重叠的三个片段,考\"该发生什么\";时间对齐维度(Alignment)的干扰是把正确片段提前、推后或拉宽,考\"何时发生\";展开维度(Progression)的干扰是倒放、乱序和循环,考\"该如何展开\"。因为题干和选项全是视频片段,选项措辞、语言先验这类文字捷径被直接焊死。默认评测每个片段抽 16 帧,一道题最多 96 帧输入。\n\n## 31 款模型的分数表:对齐维度集体塌方\n\n评测覆盖 10 款闭源和 21 款开源模型,论文直接把时间对齐称为主要瓶颈:闭源模型平均分从语义维度的 70.73%、展开维度的 67.72% 跌到对齐维度的 48.13%;开源模型平均只剩 26.54%——四选一的随机线是 25%,几乎等于乱猜。人类基线是 97%。\n\n闭源榜榜首是 Seed1.8(开思考模式),平均 88.58,对齐维度 76.92;Qwen3.5-Plus 紧随其后,平均 85.74。真正扎眼的是几款旗舰的对齐塌方:GPT5.4 平均 56.89,对齐维度只剩 37.41;Claude4.6-Sonnet 平均 51.81,对齐维度 28.80,相比自己语义维度的 55.69 几乎砍半;Grok4.1 三个维度全在 24% 上下,平均 24.11,干脆低于随机线。同厂代差也很戏剧:Seed1.6 的对齐维度只有 17.83,与 Seed1.8 隔了一代判若两款模型。\n\n开源侧,对齐维度过了 50 的只有 Qwen3.5 系两款:35B-A3B 拿到 51.55,397B-A17B 是 50.62,小杯反超大杯;平均分最高的开源模型也是 397B-A17B 的 68.27。而 Qwen3VL-32B-Instruct 的对齐分数只有 10.91,连随机线的一半都不到——「看得懂内容」和「排得出先后」在同一家人身上同时成立又同时失守。\n\n## 为什么对齐这么难:四条行为线索\n\n团队对四款代表模型做了行为敏感性分析,四条发现值得记下:候选片段换个顺序,模型的选择就不稳定;模型明显更依赖开头上下文而不是结尾;更密或更长的视觉输入反而会稀释关键时刻的边界线索;test-time scaling 能带来一些提升,但拉不平对齐维度的缺口。\n\n## 所以呢\n\n这份被 EMNLP 2026 Findings 收录的工作,值得所有拿「视频理解」当卖点的人读一遍原表。它至少提醒两件事:其一,平均分会掩盖结构性短板——GPT5.4 语义维度 68.11,对齐维度 37.41;Qwen3VL-32B-Instruct 更极端,对齐 10.91 连随机线一半都不到,但这类差距会被综合分稀释掉。其二,选型时如果业务场景是安防、剪辑、具身智能这类时序敏感方向,务必单独看对齐类指标,而不是只看综合榜。视觉模型离「真正看懂视频」,还差最后这一段时间课。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.01515","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"a137189c-9f6b-45c2-8d66-c2ff1c6412e6","en","TempCloze: open Video-LLMs near random on temporal alignment","Pick the missing middle of a video from four candidates: open Video-LLMs averaged 26.54% on temporal alignment, near random; humans scored 97%.","Video models have gotten good at recognizing what is in a frame, but a benchmark released in September by a University of Hong Kong team pours cold water on the field: cut the head and tail off a video, ask the model to pick the true missing middle from four candidate clips — a \"video cloze test\" humans pass casually — and 31 mainstream Video-LLMs get exposed ([arXiv:2609.01515](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.01515)).\n\n## How the test works\n\nTempCloze packs 1,521 videos from seven public sources, led by LVD-2M and EgoLife, favoring long-take and egocentric footage. Each video splits into a beginning (B), a missing middle (M), and an ending (E); the model gets B and E and must identify M among four candidates. The clever part is the distractors, all built from the same source video: Semantic distractors are three non-overlapping intervals, testing what should happen; Alignment distractors shift the true clip earlier, later, or wider, testing when it should happen; Progression distractors reverse, shuffle, or loop the clip, testing how it should unfold. Because prompts and answers are entirely visual, shortcuts from option wording and language priors are sealed off. Default evaluation samples 16 frames per clip, up to 96 per question.\n\n## The scoreboard: alignment collapses\n\nAcross 10 proprietary and 21 open-source models, the paper calls temporal alignment the main bottleneck: proprietary averages fall from 70.73% on Semantic and 67.72% on Progression to 48.13% on Alignment; open-source models manage just 26.54% against a 25% random line. Humans score 97%.\n\nSeed1.8 (thinking mode) tops the proprietary board at 88.58 mean with 76.92 on Alignment; Qwen3.5-Plus follows at 85.74. The flagship collapses stand out: GPT5.4 averages 56.89 with Alignment at 37.41; Claude4.6-Sonnet posts 28.80 on Alignment, nearly halving its Semantic 55.69; Grok4.1 lands at 24.11 mean, below random. The within-family gap is dramatic: Seed1.6 scores just 17.83 on Alignment, a generation behind Seed1.8.\n\nOn the open side, only two Qwen3.5 models clear 50 on Alignment — 51.55 for 35B-A3B and 50.62 for 397B-A17B, the smaller cup edging the larger — while 397B-A17B also leads open models at 68.27 mean. Qwen3VL-32B-Instruct falls to 10.91, under half the random line: understanding content and ordering events fail together in the same family.\n\n## Why alignment is hard\n\nBehavioral analysis on four representative models yields four findings: candidate reordering destabilizes choices; models lean on beginning context over ending context; denser or longer visual input dilutes decisive boundary cues; and test-time scaling brings model-dependent gains without closing the alignment gap.\n\n## So what\n\nAccepted to EMNLP 2026 Findings, this work deserves a close read from anyone selling \"video understanding.\" Two takeaways: average scores hide structural blind spots — GPT5.4 scores 68.11 on Semantic but 37.41 on Alignment, and Qwen3VL-32B-Instruct's 10.91 gets diluted in composite rankings. And for timing-sensitive deployments such as security, editing, and embodied AI, check alignment-style metrics separately instead of trusting an aggregate leaderboard. Video models still have one semester of temporal reasoning left to pass.","tempcloze-video-llm-temporal-alignment","2026-09-13T23:06:41Z","2026-09-13T23:06:44.534006Z","2026-09-13T23:06:44.534015Z",true,"agent",50,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"c3a956f8-dd42-46df-a1fd-1322dd38c15c","MentalThink 把 SVG 当作「心智草稿纸」:让多模态大模型学会用代码画心像做空间推理","mentalthink-svg-spatial-reasoning","2026-07-10T22:30:00+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"39e6e64f-a3de-4dce-bdd8-c8643f9413a1","Orca：把\"世界状态\"焊进潜空间——BAAI 推出通用世界基础模型新范式","baai-orca-world-foundation","2026-07-03T02:00:00+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"b0a2cefc-7a4e-4f2c-83a0-f1e4911f04e5","RNG-Bench：GPT-5.4\u002FGemini 3.1 Pro 闭环记忆现形","rng-bench-shanghai-ai-lab-non-markov-memory","2026-06-24T18:15:00+00:00"]