[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-opencof-wan-video-reasoning":3,"news-related-e795ec57-7401-458d-a67f-cd18098b2cf3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e795ec57-7401-458d-a67f-cd18098b2cf3","OpenCoF 把视频生成变成\"显式推理机\":字节 + 港中文用 17K 数据让 Wan 学会\"链帧思考\"","字节 Seed 与港中文 MMLab 提出 Chain-of-Frame 推理范式,基于 OpenCoF-17K(11 类任务 \u002F 17,312 条样本 \u002F 4 条策展流水线)对 Wan2.2-I2V-A14B 做 LoRA 微调,无需改架构即可在 VIPER、Gen-ViRe、MME-CoF、RULER-Bench 四个外部视频推理基准上整体超过基线;进一步引入视觉\u002F文本推理 token 机制,前者偏感知、后者偏语义,把\"视频生成模型作为纯视觉推理器\"这条新路线从论文设想推到工程可复现。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.08763v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"19cf8a89-f70f-4b77-9fa6-4f9785467d8c","en","OpenCoF turns video generation into explicit reasoning","ByteDance Seed and CUHK MMLab propose the Chain-of-Frame reasoning paradigm, based on OpenCoF-17K (11 task categories \u002F 17,312 samples \u002F 4 curation pipelines) doing LoRA fine-tuning on Wan2.2-I2V-A14B, no architecture change needed to comprehensively beat the baseline on four external video-reasoning benchmarks (VIPER, Gen-ViRe, MME-CoF, RULER-Bench); further introducing visual\u002Ftext reasoning-token mechanisms, with the former leaning toward perception and the latter toward semantics, pushing the new path of \"video generation model as a pure visual reasoner\" from paper concept to engineering reproducibility.","opencof-wan-video-reasoning","2026-07-11T18:01:00Z","2026-07-11T18:08:39.437517Z","2026-08-19T02:08:40.142862Z",true,"agent",90,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b6dc8854-6604-4860-a3de-5d70abe3e512","Real World VoiceEQ：100 万人类评分戳破语音基准饱和","hume-ai-real-world-voiceeq","2026-07-15T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"8865aca6-336a-4dbc-964a-de4afecb25c1","GenCeption 把视频生成模型改造成「通用视觉大脑」：Kaiming He 也在作者里","genception-kaiming-he","2026-07-13T10:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c3a956f8-dd42-46df-a1fd-1322dd38c15c","MentalThink 把 SVG 当作「心智草稿纸」:让多模态大模型学会用代码画心像做空间推理","mentalthink-svg-spatial-reasoning","2026-07-10T22:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"e05e3010-e356-4db8-bf15-f01c8027b937","WBench 给交互式视频世界模型做\"CT 扫描\":美团 LongCat 开源首个多轮评测基准,首测 20 个前沿模型","wbench-interactive-world-model","2026-07-05T06:01:00+00:00"]