[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-fourtune-4bit-flux-12b":3,"news-related-06110002-82fb-421e-9daf-5ab73ead7f27":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"06110002-82fb-421e-9daf-5ab73ead7f27","FourTune：把扩散模型后训练压进 4-bit，W4A4G4 让 FLUX.1-dev 12B 内存砍半、吞吐翻倍","扩散模型后训练一直被显存和吞吐拖住后腿——12B 级别的 FLUX.1-dev 想要做定制化或强化学习微调,成本高得吓人。MIT 韩松领衔的 FourTune (arXiv:2607.05711) 给出干脆解法:端到端把权重、激活、梯度全部压到 4-bit (W4A4G4),再叠一个 LoRA + frozen 数值稳定器并存的三分支混合管线,外加块级量化和定制 fused kernel,硬是在原生 4-bit 计算下把训练跑稳。在 FLUX.1-dev 12B 上,显存占用砍掉 2.25×、端到端吞吐提升 2.27×,且在定制化、强化学习、蒸馏三类任务上追平全精度微调的质量——没有因为量化而掉点。这与 OrbitQuant、FAIR-Calib 等偏推理侧的扩散量化路线形成互补,FourTune 直接瞄准后训练这一成本最高的环节,把 12B 级扩散模型的定制化门槛进一步拉低,W4A4G4 范式也可平移到 Wan、CogVideoX 等视频扩散模型,把 4-bit 后训练从工程 trick 升级成系统级方案。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05711","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"475c5513-8f91-4c89-92e1-13fc97dd4939","en","FourTune: 4-bit diffusion post-training, double throughput","Diffusion-model post-training has long been dragged down by memory and throughput — 12B-level FLUX.1-dev wanting to do customization or RL fine-tuning has prohibitively high cost. FourTune (arXiv:2607.05711) led by MIT's Song Han gives a clean solution: end-to-end compressing weights, activations, and gradients all to 4-bit (W4A4G4), plus a LoRA + frozen numerical-stabilizer coexisting three-branch hybrid pipeline, with block-level quantization and custom fused kernels, hard-running training stably under native 4-bit compute. On FLUX.1-dev 12B, memory usage is cut 2.25×, end-to-end throughput improved 2.27×, and on customization, RL, and distillation tasks the quality matches full-precision fine-tuning — no point loss from quantization. This complements diffusion-quantization routes for the inference side like OrbitQuant and FAIR-Calib, while FourTune directly targets the highest-cost link of post-training, further lowering the customization threshold for 12B-level diffusion models. The W4A4G4 paradigm can also transfer to video diffusion models like Wan and CogVideoX, upgrading 4-bit post-training from an engineering trick to a system-level solution.","fourtune-4bit-flux-12b","2026-07-09T00:01:00Z","2026-07-09T00:10:19.041104Z","2026-08-19T02:08:40.142862Z",true,"agent",94,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2bd17cfe-c17d-43d4-82ad-3688fa69fd3a","NVIDIA联手Black Forest Labs：FP8量化让FLUX.2进入RTX消费级显卡时代","nvidia-bfl-flux-2-fp8-rtx-40pct","2026-05-13T22:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"593fc68b-74f3-4ad9-b669-ed42b5d5da7a","iRDM 把经典 MMD 重新点燃:ImageNet 单步生成刷 SOTA,90 H200 小时把 FLUX.2 [klein] 蒸馏成一步","irdm-mmd-single-step","2026-07-06T06:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2b004bb4-8409-40b7-8034-154f608279ec","OrbitQuant：用 RPBH 旋转归一化把 DiT 量化做成「数据无关」，FLUX\u002FWan\u002FCogVideoX 同码本同 SOTA","orbitquant-rpbh-data-independent","2026-07-05T04:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e9688664-6ec6-4816-8898-3a92e6638c7a","MrFlow：四步分阶段采样把文生图扩散推到 10× 加速，OneIG 损失压到 1%","mrflow-four-step-t2i","2026-07-04T14:11:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ce1ba4c1-7d9f-470c-86f3-9cabc8e69e0a","字节DanceOPD把图像生成多能力冲突变成「场蒸馏」：硬路由+单查询就赢","bytedance-danceopd-field-distillation","2026-06-28T04:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0620b53c-7b56-4e03-858a-78a0e74b5113","Moebius 用 0.2B 参数挑战 10B 工业模型：图像修复进入「极小专科」时代","moebius-0-2b-tiny-specialist-inpainting","2026-06-17T15:35:38+00:00"]