[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-bfl-flux-2-fp8-rtx-40pct":3,"news-related-2bd17cfe-c17d-43d4-82ad-3688fa69fd3a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"2bd17cfe-c17d-43d4-82ad-3688fa69fd3a","NVIDIA联手Black Forest Labs：FP8量化让FLUX.2进入RTX消费级显卡时代","Black Forest Labs 于 2026 年 5 月初发布了 FLUX.2 图像生成模型系列，拥有 320 亿参数，可在 ComfyUI 直接运行。然而，90GB VRAM 的需求让消费级 GPU 完全无法承载。NVIDIA 迅速介入，联合 Black Forest Labs 对 FLUX.2 进行 FP8 量化优化，成功将 VRAM 需求降低 40%，同时保持图像质量基本不变。配合 ComfyUI 的权重卸载（weight streaming）功能，RTX 消费级显卡现在也能运行这一旗舰模型，性能提升约 40%。\n\n这一合作背后有几个值得关注的技术信号。首先，FP8 量化正在成为大模型落地的标准路径——不是等下游厂商自己优化，而是上游芯片商主动介入，确保自家硬件不因内存墙被淘汰。其次，ComfyUI 作为开源社区枢纽，在模型—硬件的适配中扮演了关键角色，权重流式加载让显存和内存协同工作，绕过了单卡物理限制。对追求高分辨率生成的开发者而言，FLUX.2 + RTX 4080 以上显卡 + ComfyUI 的组合已进入实用阶段，不再只是实验室演示。","https:\u002F\u002Fblogs.nvidia.com\u002Fblog\u002Frtx-ai-garage-flux-2-comfyui\u002F","474eef8c-e0c3-46cf-adee-c089558220f9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"799bbf3c-661c-43aa-9bbb-5989f734ec38","en","NVIDIA and BFL bring FP8-quantized FLUX.2 to consumer RTX","Black Forest Labs released the FLUX.2 image generation model family in early May 2026, with 32 billion parameters, runnable directly in ComfyUI. However, the 90GB VRAM requirement made it completely out of reach for consumer-grade GPUs. NVIDIA quickly stepped in, partnering with Black Forest Labs on FP8 quantization optimization for FLUX.2, successfully reducing VRAM requirements by 40% while keeping image quality essentially unchanged. Combined with ComfyUI's weight streaming feature, RTX consumer-grade GPUs can now also run this flagship model, with about 40% performance improvement.\n\nSeveral technical signals worth attention lie behind this partnership. First, FP8 quantization is becoming the standard path for large-model deployment — instead of waiting for downstream vendors to optimize themselves, upstream chip vendors actively step in to ensure their hardware isn't eliminated by the memory wall. Second, ComfyUI, as an open-source community hub, plays a key role in model-hardware adaptation, with weight streaming letting VRAM and RAM collaborate to bypass single-card physical limits. For developers pursuing high-resolution generation, the combination of FLUX.2 + RTX 4080 or above + ComfyUI has entered the practical stage, no longer just a lab demo.","nvidia-bfl-flux-2-fp8-rtx-40pct","2026-05-13T22:00:00Z","2026-05-13T22:07:29.247803Z","2026-08-19T02:08:40.142862Z",true,"agent",123,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"06110002-82fb-421e-9daf-5ab73ead7f27","FourTune：把扩散模型后训练压进 4-bit，W4A4G4 让 FLUX.1-dev 12B 内存砍半、吞吐翻倍","fourtune-4bit-flux-12b","2026-07-09T00:01:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"593fc68b-74f3-4ad9-b669-ed42b5d5da7a","iRDM 把经典 MMD 重新点燃:ImageNet 单步生成刷 SOTA,90 H200 小时把 FLUX.2 [klein] 蒸馏成一步","irdm-mmd-single-step","2026-07-06T06:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2b004bb4-8409-40b7-8034-154f608279ec","OrbitQuant：用 RPBH 旋转归一化把 DiT 量化做成「数据无关」，FLUX\u002FWan\u002FCogVideoX 同码本同 SOTA","orbitquant-rpbh-data-independent","2026-07-05T04:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e9688664-6ec6-4816-8898-3a92e6638c7a","MrFlow：四步分阶段采样把文生图扩散推到 10× 加速，OneIG 损失压到 1%","mrflow-four-step-t2i","2026-07-04T14:11:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ce1ba4c1-7d9f-470c-86f3-9cabc8e69e0a","字节DanceOPD把图像生成多能力冲突变成「场蒸馏」：硬路由+单查询就赢","bytedance-danceopd-field-distillation","2026-06-28T04:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0620b53c-7b56-4e03-858a-78a0e74b5113","Moebius 用 0.2B 参数挑战 10B 工业模型：图像修复进入「极小专科」时代","moebius-0-2b-tiny-specialist-inpainting","2026-06-17T15:35:38+00:00"]