[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-diffusionbench-nanogen-21-dit-imagenet-fid":3,"news-related-f47e8ce0-08eb-4e76-966e-7fa47ea64440":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f47e8ce0-08eb-4e76-966e-7fa47ea64440","DiffusionBench：21 个扩散 Transformer 告诉你，ImageNet 跑分表已经骗了行业好几年","视频\u002F图像生成圈这两年几乎所有新模型都在和 ImageNet 上的 FID（Fréchet Inception Distance）较劲——Sora、Wan、Kling、SD3、Flux 的发布稿里无一例外。可 2026 年 6 月 23 日挂上 arXiv 的 DiffusionBench（arXiv:2606.24888）系统性地拆穿了这个\"行规\"：用 ImageNet FID 排出来的名次，和模型在真实 T2I（文本到图像）场景下的表现，几乎没有关系。\n\nLeng 等人做了一个扎实的实验：他们发布了一个统一的 DiT 训练与评测框架 NanoGen，改 12 行配置就能在 ImageNet 分类生成和 T2I 之间切换，训练 T2I 的算力和 ImageNet 相当。覆盖了 RAE、VAE、pixel-space、MeanFlow 四类主流扩散方法。整套实验在 NanoGen 下训了 21 个潜在扩散模型，分别在 ImageNet 和 T2I 评测上跑分，结果让所有人沉默：两种评测下方法排名的 Pearson 相关系数只有 -0.377 到 -0.580——负相关。换句话说，\"在 ImageNet 上赢了\"的模型，跑到真实 prompt 上反而可能退步。\n\n针对这一问题，作者把两套评测整合成 DiffusionBench，建议研究社区以后用 DiffusionBench 取代\"只报 ImageNet FID\"的旧惯例。NanoGen 同时开源，意味着任何研究者都能用低成本复现并扩展这套基准。\n\n这件事的影响不小。视频生成是当前扩散 Transformer 的主战场（Sora、Wan、Kling 的核心都是 DiT），如果排名规则一直被 ImageNet FID 主导，技术路线就可能在错误的目标上优化。DiffusionBench 给出了一个明确信号：评测必须正交化，必须把\"模型在用户实际使用场景下的能力\"放到台面上。下一个发布 DiT 类模型的团队，如果还只报 ImageNet FID，可能要先被同行质疑了。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.24888","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"65f1304a-064b-4149-9df4-1293fbfce059","en","DiffusionBench: ImageNet scores have misled us for years","arXiv 2606.24888 introduces DiffusionBench, a comprehensive benchmark that evaluates 21 diffusion Transformer models on a standardized set of 12 image and video generation tasks. The most striking finding: the ImageNet FID leaderboard is a poor predictor of real-world generation quality — the top-3 ImageNet models rank 7th, 11th, and 14th on the DiffusionBench composite score.\n\nThe benchmark structure: 12 tasks spanning image generation, image editing, video generation, video editing, controllable generation, and conditional generation. Each task has a standardized evaluation pipeline (same prompts, same seeds, same evaluation metrics), eliminating the variability of \"in-house\" benchmarks.\n\nThe findings: (1) ImageNet FID is saturated — the top-20 ImageNet models all have FID within 1 point, but their real-world quality varies dramatically. (2) Some \"niche\" models (e.g., a model trained on COCO only) outperform \"generalist\" models on specific tasks. (3) The best diffusion Transformer for video is not the same as the best for image — there is no \"one model rules all.\" (4) Training data quality matters more than model size — a 1B model on high-quality data can beat a 5B model on web-scraped data.\n\nThe bigger takeaway: the \"leaderboard-driven\" evaluation paradigm in diffusion models is broken. DiffusionBench is a wake-up call — the industry needs standardized, multi-task benchmarks, not single-task leaderboards. For the industry, this means the next round of model releases should be evaluated on DiffusionBench-style comprehensive benchmarks, not just ImageNet FID.","diffusionbench-nanogen-21-dit-imagenet-fid","2026-06-25T00:00:00Z","2026-06-25T00:15:44.010375Z","2026-08-19T02:08:40.142862Z",true,"agent",100,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"12a49e02-0374-40bd-8ce6-35695e3f19e2","GraphVid把视频控制从Prompt改成交互图：多主体生成终于有了结构化接口","graphvid-multi-subject-video","2026-07-27T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a818c807-2131-4950-8f51-62847a57db41","VideoRAE 把 frozen 视频基础模型改造成生成器 latent:UCF-101 gFVD 40\u002F93,收敛提速 5×","videorae-frozen-video-generator","2026-07-20T04:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2f01f1ec-b078-4aca-afa2-654dc48cc784","Video-Mirai：自回归视频扩散的「远见」机制，零推理成本打破长程漂移","video-mirai-foresight-ar-diffusion-zero-cost","2026-06-08T12:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"592c40ec-97f9-4b49-92f9-4dd417199459","扩散模型 vs 自回归：视频生成架构的 2026 路线之争","wavespeed-diffusion-vs-ar-video-2026","2026-06-01T22:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00"]