[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-scail-2-zhipu-tsinghua-end-to-end-skeleton":3,"news-related-65c90f98-0017-45ff-869e-a8cd251d7498":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"65c90f98-0017-45ff-869e-a8cd251d7498","SCAIL-2：智谱+清华用端到端架构改写角色动画\"骨架法则\"","视频生成模型过去一年在画质、长度上飞速进化，但角色动画这条赛道始终跳不出\"先抽骨架再驱动\"的旧范式。智谱AI 与清华刘永进教授课题组联合开源的 SCAIL-2，把这条范式切开了。\n\nSCAIL-2 的核心突破在于彻底放弃 2D 关键点与 SMPL Mesh 等显式中间表示，直接在像素级拼接驱动视频的隐空间特征与参考角色特征，让模型用\"视觉直觉\"而非\"符号翻译\"理解运动。配合 DiT 架构中的全上下文姿态注入和 Pose-Shifted RoPE，模型在多人复杂交互、动物驱动零样本泛化等传统方案几乎失灵的场景里跑通。SCAIL-2 支持 512p\u002F704p 双分辨率，Apache 2.0 协议开源，权重同步上架 Hugging Face、ModelScope 和 GitHub，ComfyUI 工作流开箱即用。\n\n更深层的工程意义是端到端带来的算力简化：传统管线需要骨架提取、姿态重投影、掩码生成多个串行环节，SCAIL-2 全部塞进一个 Transformer，推理延迟与显存占用显著下降。智谱构建的\"AI 合成 AI 数据\"工厂化管线，让角色动作从\"火柴人\"演变为可复用视觉向量，对游戏、直播、影视数字人产业链具备直接商业价值。\n\nSCAIL-2 仍有边界：手部、面部等细颗粒度控制仍依赖大规模高质量配对数据。但\"工业级精准控制\"这条路线，比单纯卷参数量的视频模型更接近真正的生产工具需求，也是 2026 年视频生成走向产业化的关键信号。","https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FSCAIL-2","1eab5c4a-0c8e-49a4-8ac8-0f84a2a3c3a4",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c70aa504-e694-493e-a929-47437d8fe465","en","SCAIL-2: Zhipu and Tsinghua rewrite character animation","Video generation models have advanced rapidly in quality and length over the past year, but character animation has remained stuck in the old \"extract skeleton first, then drive\" paradigm. Zhipu AI and Professor Liu Yongjin's group at Tsinghua jointly open-sourced SCAIL-2, cutting that paradigm open.\n\nSCAIL-2's core breakthrough is the complete abandonment of explicit intermediate representations like 2D keypoints and SMPL Mesh, directly stitching together the latent-space features of the driving video and the reference character at the pixel level, letting the model understand motion through \"visual intuition\" rather than \"symbolic translation.\" Combined with full-context pose injection in the DiT architecture and Pose-Shifted RoPE, the model handles scenarios where traditional approaches nearly fail — multi-person complex interactions, zero-shot animal driving generalization. SCAIL-2 supports 512p\u002F704p dual resolution, ships under Apache 2.0, with weights on Hugging Face, ModelScope, and GitHub, and ComfyUI workflows ready out of the box.\n\nThe deeper engineering significance is the compute simplification that comes with end-to-end: traditional pipelines require multiple serial steps — skeleton extraction, pose re-projection, mask generation — and SCAIL-2 folds them all into a single Transformer, significantly reducing inference latency and VRAM. Zhipu's \"AI synthesizes AI data\" factory-style pipeline is letting character motion evolve from \"stick figures\" to reusable visual vectors, with direct commercial value for games, livestreaming, and film\u002FTV digital-human production lines.\n\nSCAIL-2 still has its limits: fine-grained control of hands and faces still depends on large-scale high-quality paired data. But the \"industrial-grade precise control\" path is much closer to actual production-tool needs than simply scaling parameters — and it's a key signal that video generation is industrializing in 2026.","scail-2-zhipu-tsinghua-end-to-end-skeleton","2026-06-10T06:00:00Z","2026-06-11T14:13:18.092717Z","2026-08-19T02:08:40.142862Z",true,"agent",120,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"aad00b18-d354-48b5-ad21-62b53150b8c6","MiniMax H3 开源实测:你下载的权重,和 API 里跑的不是同一个模型","minimax-h3-local-vs-api-gap","2026-08-15T17:07:24+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00"]