[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-videochat3-4b-mllm":3,"news-related-9f82c248-0592-421f-9fd8-ebd2100dcaf5":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9f82c248-0592-421f-9fd8-ebd2100dcaf5","VideoChat3 全开源 4B 视频 MLLM 一次打通四种能力,I3D-ViT 把时空 token 砍掉 16×","南京大学 MCG 团队（S-Lab）近日开源 VideoChat3——4B 参数全开放视频 MLLM，登顶 Hugging Face 趋势榜前 3，主打「一个模型搞定所有视频理解」：细粒度运动感知、小时级长视频、时序 grounding、流式响应四能力合一。\n\n两个核心创新：\n- **I3D-ViT**：时空 token 做 16× 压缩，显著降低视觉编码成本。\n- **Adaptive Frame Resolution for Streaming**：按需提升帧分辨率，避免每帧高分辨率计算。\n\n团队同步开源三套数据集——Academic2M（通用）、LV116K（长视频）、OL617K（流式），覆盖训练全链路。VideoChat3 把模型权重、数据、合成 pipeline 一次性 deliver——这在视频 MLLM 圈相当罕见，多数强模型只放权重、不放数据。\n\n实验显示 VideoChat3 在通用、长视频、流式三类基准上均超过参数规模更大的开源对手。统一架构 + 自适应帧率 + 紧凑 token 表示，把四种能力塞进 4B——正是视频 MLLM 走向工程化的关键一步。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.14935","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"0733e1e4-77f2-4e9c-bc30-fd6c4fc89462","en","VideoChat3: open 4B video MLLM, 16x fewer tokens","Nanjing University's MCG team (S-Lab) recently open-sourced VideoChat3 — a 4B-parameter fully open video MLLM that topped the Hugging Face trending top 3, with \"one model handling all video understanding\" as the main feature: fine-grained motion perception, hour-level long video, temporal grounding, and streaming response all in one. Two core innovations: **I3D-ViT**: 16× compression of spatiotemporal tokens, significantly reducing visual encoding cost. **Adaptive Frame Resolution for Streaming**: on-demand frame resolution upscaling, avoiding per-frame high-resolution computation. The team simultaneously open-sourced three datasets — Academic2M (general), LV116K (long video), OL617K (streaming) — covering the entire training chain. VideoChat3 delivers model weights, data, and synthesis pipeline in one go — this is quite rare in the video MLLM circle, where most strong models only release weights, not data. Experiments show VideoChat3 outperforms open-source opponents with larger parameter scales on all three categories: general, long video, and streaming. Unified architecture + adaptive frame rate + compact token representation, packing four capabilities into 4B — that's the key step for video MLLM to move toward engineering.","videochat3-4b-mllm","2026-07-15T02:00:00Z","2026-07-17T18:12:26.125040Z","2026-08-19T02:08:40.142862Z",true,"agent",171,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"3851a096-37d6-4a45-bfb7-57b2fd65d992","京东开源 JoyAI-Echo：5 分钟长视频生成首次解决「跨镜头一致性」难题，DMD 蒸馏跑出 7.5× 加速","joyai-echo-jd-5-min-cross-shot-dmd-7-5x","2026-06-12T02:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00"]