[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-rhymeflow-tsinghua-async-denoising-1-53x":3,"news-related-42cfc778-8f1b-4bf2-a0ae-4343a066f48d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"42cfc778-8f1b-4bf2-a0ae-4343a066f48d","RhymeFlow：清华提出异步去噪流调度，DiT视频生成训练免费加速1.53倍","【核心思路】清华大学与GigaAI联合发布RhymeFlow框架，提出异步去噪流调度（Asynchronous Denoising Flow Scheduling）机制，无需重训练即可显著加速基于DiT（Diffusion Transformer）的视频生成模型。论文于6月4日上线arXiv（2606.06309），代码以Apache-2.0协议开源。\n\n【技术突破】现有训练免费加速方法（如SVG、SAP、DiCache）多聚焦于\"单个去噪步内的注意力稀疏化\"，但仍然要求视频中每一帧在全部时间步上完成完整的密集去噪。RhymeFlow打破这一刚性约束，将视频帧分成\"关键帧\"与\"非关键帧\"两类：关键帧锚定语义转换，保留密集逐步去噪以保结构完整；非关键帧按\"节奏感\"渐进跳过可预测的去噪步，仅通过轻量\"潜空间轨迹投影\"在3D注意力中维持时序一致性。\n\n【性能数据】在Wan 2.1上RhymeFlow以PSNR 26.29、SSIM 0.783超越SAP（24.45\u002F0.730），实现1.53倍加速；与SAP组合后速度达1.66倍。在HunyuanVideo上，单独使用实现2.26倍加速，叠加SAP更达到2.60倍的极致加速，且视觉质量（PSNR\u002FSSIM\u002FLPIPS）全面优于SVG、EasyCache、DiCache、VGDFR等基线。\n\n【观点】RhymeFlow体现了一种\"正交加速维度\"——不改变模型权重，只重新组织推理时序。这与稀疏注意力、KV-Cache、投机解码方向高度互补。对于DiT视频模型这种\"算力巨兽\"而言，\"调度即优化\"的思路，可能是2026年下半年推理成本继续下探的最务实路径之一。\n\n【出处】arXiv: 2606.06309（2026-06-04）；GitHub: Simon-Dcs\u002FRhymeFlow","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.06309","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"6f05a702-2298-4cf3-92d1-e7bada9e3421","en","RhymeFlow: async denoising flow gives DiT training free 1.53x","**Core idea.** Tsinghua University and GigaAI jointly released the RhymeFlow framework, proposing an asynchronous denoising flow scheduling mechanism that significantly accelerates DiT (Diffusion Transformer)-based video generation models without retraining. The paper went online on arXiv on June 4 (2606.06309), with code open-sourced under Apache 2.0.\n\n**Technical breakthrough.** Existing training-free acceleration methods (such as SVG, SAP, DiCache) mostly focus on \"attention sparsification within a single denoising step,\" but still require every frame in the video to complete full dense denoising across all timesteps. RhymeFlow breaks this rigid constraint, dividing video frames into \"key frames\" and \"non-key frames\": key frames anchor semantic transitions, preserving dense stepwise denoising to keep structure intact; non-key frames skip predictable denoising steps progressively according to a \"rhythmic sense,\" maintaining temporal consistency only through lightweight \"latent-space trajectory projection\" within 3D attention.\n\n**Performance data.** On Wan 2.1, RhymeFlow surpasses SAP (24.45\u002F0.730) with PSNR 26.29 and SSIM 0.783, achieving a 1.53× speedup; combined with SAP, the speed reaches 1.66×. On HunyuanVideo, used alone it achieves a 2.26× speedup, and stacked with SAP it reaches a staggering 2.60× speedup, with visual quality (PSNR\u002FSSIM\u002FLPIPS) comprehensively superior to baselines like SVG, EasyCache, DiCache, and VGDFR.\n\n**Perspective.** RhymeFlow embodies an \"orthogonal acceleration dimension\" — not changing model weights, but reorganizing the inference timeline. This is highly complementary to sparse attention, KV-Cache, and speculative decoding. For DiT video models, the \"compute behemoths,\" the \"scheduling-is-optimization\" line of thinking may be one of the most pragmatic paths for inference cost to keep falling in the second half of 2026.\n\n**Source:** arXiv: 2606.06309 (2026-06-04); GitHub: Simon-Dcs\u002FRhymeFlow","rhymeflow-tsinghua-async-denoising-1-53x","2026-06-07T22:00:00Z","2026-06-07T22:13:34.943819Z","2026-08-19T02:08:40.142862Z",true,"agent",173,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"0599b775-ac17-49d2-aebd-a16f531c7168","腾讯混元 MeanFlowNFT：把 RL 接进「平均速度生成器」，Wan 2.1 4 步反超 50 步 LongCat-Video RL","tencent-hunyuan-meanflownft","2026-07-16T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bf8755fc-cd4f-4bd9-9617-e70f56ddc4ac","LTX-2.3：开源视频生成正式进入 4K + 原生音频时代","ltx-2-3-lightricks-4k-native-audio","2026-06-02T01:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5612d186-46ee-4509-9a93-94045ba004ae","LTX-2.5 开放权重视频模型:4K 反而在 Fast 端点,EXR 色彩管线也焊进去了","ltx-2-5-open-weights-video","2026-08-18T15:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00"]