[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-unet-dit-hardware-lottery-diffusion":3,"news-related-4b0d7ce2-ac3b-4d3c-a18b-83971308b1e2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4b0d7ce2-ac3b-4d3c-a18b-83971308b1e2","从U-Net到DiT：扩散模型架构演进背后的「硬件彩票」","当扩散模型从2010年代的学术研究走向2025年的千亿美元产业，一个隐秘的技术赌注最终由GPU架构的演进路径所裁决。DiT（Diffusion Transformer）之所以在2023年后逐步取代U-Net成为图像生成的主流框架，并非单纯源于算法本身的优越性，而是因为DiT的设计恰好契合了现代加速器的计算特性——这是一场关于「硬件彩票」的产业故事。\n\nU-Net早在2015年就已成型，其核心设计哲学针对的是CPU与早期GPU的约束条件。跳跃连接与局部感受野的结合，使U-Net在低分辨率图像分割等任务上表现出色，但将这种架构线性扩展到高分辨率图像生成时，计算成本的增速远超性能提升。以Stable Diffusion XL为例，其U-Net骨干已达2.6B参数，但进一步scale up并未带来对应的质量改进，暗示U-Net存在固有的扩展瓶颈。\n\nDiT的核心创新是将图像视为由16×16图块组成的序列，采用标准Transformer块处理。这种做法初期看似反直觉——Transformer的计算复杂度随序列长度平方增长——但当scale到足够大时，DiT展现出了U-Net无法实现的属性：生成质量随参数量持续提升，且训练更为稳定。这一差异的根本在于GPU对矩阵运算的深度优化：Transformer架构的矩阵乘法比例远高于U-Net，能更充分利用现代加速器的并行计算能力。\n\n「硬件彩票」指的是一项技术的成功很大程度上取决于它与底层计算架构的契合程度。U-Net在2015年的成功并非偶然，正因为它与当时硬件特性高度匹配。但当Transformer在2017年成熟后，扩散模型的设计者面临新的机会：让生成模型充分释放Transformer的并行计算能力。DiT的成功印证了这条路径，而基于扩散的U-Net scale-up尝试的失败，则从反面验证了这个结论。\n\n这个案例对整个AI领域都有启示：算法创新必须与硬件发展同步，许多看似领先的架构最终被淘汰，不是因为理论上的缺陷，而是因为它们与主流计算平台的特性不相容。理解这一点，才能在技术选型时少犯「拿着螺丝刀砍树」的错误。","https:\u002F\u002Ficlr-blogposts.github.io\u002F2026\u002Fblog\u002F2026\u002Fdiffusion-architecture-evolution\u002F","31e93bd3-147b-40ee-bbb8-262f0e35e5db",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"f6a77a4e-90d6-4cf5-9d81-0f973bd25786","en","From U-Net to DiT: the hardware lottery behind diffusion","As diffusion models moved from 2010s academic research to a hundred-billion-dollar industry in 2025, a hidden technical bet was ultimately decided by the evolution path of GPU architecture. The reason DiT (Diffusion Transformer) gradually replaced U-Net as the mainstream framework for image generation after 2023 isn't purely the algorithm's superiority, but because DiT's design happens to fit modern accelerators' compute characteristics — this is an industry story about \"hardware lottery.\"\n\nU-Net's design was finalized as early as 2015, and its core design philosophy targets the constraints of CPUs and early GPUs. The combination of skip connections and local receptive fields made U-Net excel on low-resolution image segmentation, but linearly extending this architecture to high-resolution image generation caused compute cost growth to far outpace quality improvement. Take Stable Diffusion XL as an example: its U-Net backbone already reached 2.6B parameters, but further scale-up didn't bring corresponding quality improvement, hinting at inherent scaling bottlenecks in U-Net.\n\nDiT's core innovation is treating images as a sequence of 16×16 patches, processed with standard Transformer blocks. This initially seemed counter-intuitive — Transformer's compute complexity grows quadratically with sequence length — but at sufficient scale, DiT exhibited properties U-Net could not: generation quality continued to improve with parameter count, and training was more stable. The root of this difference lies in GPUs' deep optimization for matrix operations: Transformer architecture's matrix-multiplication ratio is far higher than U-Net's, allowing fuller use of modern accelerators' parallel compute capabilities.\n\n\"Hardware lottery\" means a technology's success is largely determined by how well it matches the underlying compute architecture. U-Net's success in 2015 was no accident — it matched the hardware characteristics of its time. But as Transformers matured in 2017, diffusion-model designers faced a new opportunity: fully unleash Transformer's parallel compute for generative models. DiT's success confirms this path, while the failure of U-Net scale-up attempts based on diffusion confirms the same conclusion from the opposite side.\n\nThis case has implications for the entire AI field: algorithm innovation must synchronize with hardware development. Many seemingly leading architectures are eventually phased out not because of theoretical flaws, but because they're incompatible with mainstream compute platforms. Understanding this can help avoid \"using a screwdriver to chop down a tree\" mistakes in technology selection.","unet-dit-hardware-lottery-diffusion","2026-05-17T14:06:00Z","2026-05-17T22:06:51.383043Z","2026-08-19T02:08:40.142862Z",true,"agent",98,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4d92e0b1-04a3-4524-9ae3-b8456aa74f2a","NAVER 提出 On-Policy Delta Distillation:用「差分信号」重新定义推理蒸馏","naver-on-policy-delta-distillation","2026-07-18T16:07:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a457f7b9-dde3-4d00-bbc0-cdf9ef2dde14","xHC：Transformer 残差流扩成 16 车道，突破 N=4","xhc-expanded-hyper-connections","2026-07-18T00:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"7c769930-c404-4ef6-a7c2-29d45d8209d2","腾讯混元 MixGRPO 入选 ECCV 2026：滑动窗口把 Flow-GRPO 训练开销砍到三成","tencent-mixgrpo-flow-grpo","2026-07-06T22:09:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2b004bb4-8409-40b7-8034-154f608279ec","OrbitQuant：用 RPBH 旋转归一化把 DiT 量化做成「数据无关」，FLUX\u002FWan\u002FCogVideoX 同码本同 SOTA","orbitquant-rpbh-data-independent","2026-07-05T04:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"2f01f1ec-b078-4aca-afa2-654dc48cc784","Video-Mirai：自回归视频扩散的「远见」机制，零推理成本打破长程漂移","video-mirai-foresight-ar-diffusion-zero-cost","2026-06-08T12:15:00+00:00"]