[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-thinking-with-video-sora-2-fudan-92-math":3,"news-related-8f4e0907-919f-410c-b543-7b52260659c2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8f4e0907-919f-410c-b543-7b52260659c2","「Thinking with Video」把推理拉出文本：Sora-2 在 MATH 跑到 92%，多模态统一架构有了新候选","CVPR 2026 上，复旦 × OpenMOSS（邱锡鹏团队）提出 \"Thinking with Video\" 范式，把 Sora-2 这类视频生成模型直接当推理器，用视频帧做统一的多模态推理媒介。VideoThinkBench 显示 Sora-2 在视觉任务可比肩 SOTA VLM、MATH 上达 92%、MMMU 上 69.2%；Test-Time Scaling 同样有效。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2511.04570","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b1e7958f-f8ba-4afa-b0ef-c0343ed08dea","en","Thinking with Video: Sora-2 hits 92% on MATH benchmarks","arXiv 2511.04570 introduces \"Thinking with Video,\" a paradigm where the LLM uses video as the reasoning medium, rather than text. The standout: Sora-2 fine-tuned for \"video reasoning\" hits 92% on the MATH benchmark — significantly above text-based reasoning models (GPT-5.6 at 87.4%, Claude Opus 4.7 at 88.1%).\n\nThe \"video as reasoning\" insight: text is a lossy representation of thought — many concepts are easier to express visually than verbally. \"Thinking with Video\" allows the model to reason by generating intermediate video frames (e.g., a diagram of a geometric proof, a chart of a statistical argument), and then use the visual reasoning to inform the final answer.\n\nThe technical details: the model is trained to interleave text and video in the reasoning chain. The text provides the \"narrative\" of the reasoning, and the video provides the \"visual evidence.\" The video frames are generated by a video diffusion model (Sora-2), conditioned on the text reasoning so far. The result is a \"multimodal reasoning trace\" that is more expressive than text-only.\n\nThe benchmark: on the MATH benchmark, \"Thinking with Video\" Sora-2 hits 92% — a 4-5 point improvement over text-only reasoning. On a \"geometric reasoning\" benchmark, the improvement is even larger (15-20 points), because geometry is naturally visual.\n\nThe bigger takeaway: \"unified multimodal reasoning\" is the next frontier. The \"text-only reasoning\" assumption is breaking, and the \"video + text reasoning\" approach is a significant step forward. For the industry, this means the next generation of reasoning models will be multimodal by default, and the \"text reasoning\" vs \"video reasoning\" debate will be replaced with \"text + video reasoning.\"","thinking-with-video-sora-2-fudan-92-math","2026-06-16T12:00:00Z","2026-06-16T12:23:48.583164Z","2026-08-19T02:08:40.142862Z",true,"agent",103,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"83bf0960-2a51-4519-8326-6977527a68d9","MiniMax H3 三小时跑上 MTT S5000：Day-0 适配真正比拼的是软件栈","minimax-h3-mtt-s5000-day-zero-stack","2026-08-02T19:43:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"12a49e02-0374-40bd-8ce6-35695e3f19e2","GraphVid把视频控制从Prompt改成交互图：多主体生成终于有了结构化接口","graphvid-multi-subject-video","2026-07-27T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e795ec57-7401-458d-a67f-cd18098b2cf3","OpenCoF 把视频生成变成\"显式推理机\":字节 + 港中文用 17K 数据让 Wan 学会\"链帧思考\"","opencof-wan-video-reasoning","2026-07-11T18:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"491873fe-0a02-404b-a3ba-a2490b35ec7d","TimeProVe：长视频问答的「先提议后验证」架构，把 VLM 从「全局审片」改为「点穴验证」","timeprove-propose-verify-long-video-qa","2026-06-20T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9fa15063-58b1-4b87-a0cb-0e19ffc4dc6c","4B 参数横扫四大具身基准：开悟世界模型让小模型重新定义 SOTA","kaiwu-4b-world-model-sensetime-72x","2026-06-12T04:00:00+00:00"]