[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-flux-3-image-bounding-boxes":3,"topics-all":38,"news-related-473ca44d-b179-4c19-a7a6-916dd8b90e4a":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"473ca44d-b179-4c19-a7a6-916dd8b90e4a","FLUX 3 Image 发布:画框控图,逐框编辑","Black Forest Labs 发布 FLUX 3 家族图像组件 FLUX 3 Image:画布按 0-1000 网格,元素用边界框声明位置与内容,模型逐框生成;编辑只动目标框,其余像素不变。布局提示为全局说明加 JSON 元素表,可交由 LLM 规划。支持 10 张参考图、原生 4K 与商业权重自托管。","AI 生图这一两年最持久的抱怨不是画质,而是可控性:提示词写三行,布局靠运气,想改画面一角,整张图重掷一次骰子。10 月 1 日,Black Forest Labs 发布 FLUX 3 家族的图像组件 FLUX 3 Image,把「控制」做成了产品的一等公民——你先在画布上画框,模型负责把每个框里的内容画对,框外的像素一个都不动。\n\n## 画框即指令:结构化布局进 prompt\n\nFLUX 3 是 BFL 的多模态模型家族,覆盖视频、音频、图像与动作,FLUX 3 Image 负责其中的图像生成与编辑。它的核心交互是 bounding box:画布在横竖两个方向都按 0 到 1000 的网格划分,每个元素用一个边界框声明,格式是 [y_min, x_min, y_max, x_max],配一段文字描述。\n\n完整的布局提示分两部分:一段全局 caption 描述整张图,后面跟一张 JSON 元素表,每行包含 id、边界框和描述;caption 里用 id 引用每个元素(比如 animal_1 首次出现的位置)。这本质上是把「自然语言生图」升级成「结构化声明生图」——位置、内容、关系都显式写在表里,不再靠模型从一句话里猜构图。\n\n不想自己画框?官方的方案是让 LLM 代劳:给一行描述和宽高比,由 LLM 规划出 caption 加元素表,每个框生成后仍可单独移动调整。模型侧还有一个 prompt upsampler,把短请求扩写成训练用的 dense caption,但你画的每个框会原样送达模型,id 和坐标都不变。\n\n## 逐框编辑:改一处,不动其他\n\n编辑能力同样围绕框展开:每个框可以重新描述、替换成别的东西或移动位置,没碰到的部分保持原样。官方页把它叫 pixel-perfect edit——加两个潜水员、换一只鸭子、给衬衫印图案,一轮一轮改下去,图不会散架。这个「局部修改、全局不变」的保证,正是过去 AI 生图在生产管线里最被诟病的短板,TechTimes 的报道直接把它归因为 LLM 驱动流水线此前缺失的能力。\n\n其他规格:单次可组合最多 10 张参考图;原生按全分辨率渲染,官方示例给出 5456 × 3072 像素的 4K 输出;企业可以走商业权重许可,在自己的基础设施上微调和部署。\n\n## 为什么值得注意:这是给 Agent 留的接口\n\n单看是「画框生图」的功能更新,放进更大的语境里,FLUX 3 Image 的接口设计明显在为 Agent 时代铺路。JSON 进 JSON 出、坐标可验证、改动可局部化——这些恰好是 LLM 管线消费图像模型时需要的属性:上游模型规划布局,下游模型按框验收,每一步都可检查、可回滚。相比之下,纯自然语言提示词接口对 Agent 来说是黑盒,「改对了没有」只能靠再套一层视觉模型去猜。\n\nBFL 自己也把这层意思写在了页面里:一个 agent 只需一条文本提示,就能用模型创建构图良好的图像。配合商业权重许可策略,FLUX 3 Image 想吃的显然不只是设计师工具市场,还有把生图能力嵌进自动化工作流的企业订单。\n\n对普通用户的「所以呢」是:生图正在从「碰运气的许愿池」变成「可规划的画布」。当位置可以声明、修改可以局部化,图像生成的使用方式会越来越像排版和工程,而不是抽卡——下一次你抱怨 AI 改图改不对,不妨想想:问题是不是出在你还在用一句话,而接口已经进化到表格了。\n\n参考:官方模型页 bfl.ai\u002Fmodels\u002Fflux-3-image ;The Decoder the-decoder.com(black-forest-labs-launches-flux-3-image)","https:\u002F\u002Fbfl.ai\u002Fmodels\u002Fflux-3-image","12897aab-bc2f-4ce3-9a8d-8be683b675ef",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"9b89cef8-30aa-4fd5-8706-15d5b8b1a9fd","en","FLUX 3 Image: bounding boxes, pixel-perfect edits, 4K","BFL's FLUX 3 Image: 0-1000 grid, caption plus JSON element tables, box-by-box edits leaving other pixels intact, 10 references, native 4K output.","Image generation's longest-running complaint is not quality but control: you write three lines of prompt, composition is a coin flip, and touching one corner means re-rolling the whole picture. On October 1, Black Forest Labs released FLUX 3 Image, the image component of its multimodal FLUX 3 family, and it makes control a first-class citizen — you draw the boxes, the model paints inside them, and every pixel outside stays put.\n\n## Bounding boxes as instructions\n\nFLUX 3 is BFL's multimodal family covering video, audio, images and actions; FLUX 3 Image handles image generation and editing. The core interaction is the bounding box. The canvas is a 0-to-1000 grid on both axes, and each element is declared with a box in [y_min, x_min, y_max, x_max] format plus a text description.\n\nA full layout prompt has two parts: a global caption describing the whole image, followed by a JSON element table where each row carries an id, a bounding box and a description. The caption cites every element by its id — animal_1 and friends — at first mention. This upgrades natural-language generation into structured declaration: position, content and relationships sit explicitly in the table, not guessed from one sentence.\n\nYou don't have to draw the boxes yourself. Hand an LLM one line and an aspect ratio, and it plans the caption and element table; every box stays editable afterward. On the model side, a prompt upsampler expands short requests into the dense captions FLUX 3 was trained on, but every box you drew reaches the model verbatim, with the same id and coordinates.\n\n## Edit one box, keep the rest\n\nEditing works box by box: re-describe a box, swap its content, or move it — everything untouched stays exactly where it was. BFL calls it pixel-perfect editing: add two divers, replace the duck, print on a shirt, round after round without the image falling apart. TechTimes frames this exact guarantee as the capability gap that has kept AI image editing unreliable in production workflows and LLM-driven pipelines.\n\nOther specs: up to 10 reference images per composition, native full-resolution rendering (an official example shows 5456 × 3072 output), and a commercial weights license that lets companies fine-tune and deploy on their own infrastructure.\n\n## An interface built for agents\n\nSeen as a feature update, this is nice; seen as interface design, it is clearly aimed at the agent era. JSON in, JSON out, verifiable coordinates, localized edits — exactly what an LLM pipeline needs when it consumes an image model: an upstream model plans the layout, a downstream model verifies box by box, and every step is checkable and reversible. Plain-prompt interfaces are black boxes by comparison; whether an edit landed correctly needs another vision model to guess.\n\nBFL says it outright on the page: an agent can use the model to create well-composed images with just a text prompt. Pair that with the commercial-weights licensing strategy and the target market is clearly bigger than designer tools — it is companies wiring image generation into automated workflows.\n\nThe takeaway: image generation is shifting from a wish-granting well to a plannable canvas. When position can be declared and edits localized, using image models starts to look like layout and engineering rather than gacha. Next time AI cannot edit your picture right, consider whether the problem is that you are still using one sentence — while the interface moved on to tables.\n\nReferences: official model page https:\u002F\u002Fbfl.ai\u002Fmodels\u002Fflux-3-image ; The Decoder https:\u002F\u002Fthe-decoder.com\u002Fblack-forest-labs-launches-flux-3-image-with-multi-step-editing-that-leaves-the-rest-of-your-picture-alone\u002F","flux-3-image-bounding-boxes","2026-10-02T23:06:31Z","2026-10-02T23:07:45.332994Z","2026-10-02T23:07:45.333007Z",true,"agent",538,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"59d14de3-0e17-4202-8e7d-ad0bc51e3471","Qwen-Image-2.1-Turbo开源:8步去噪出图","qwen-image-2-1-turbo-8-step","2026-10-11T15:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"42d37ae0-3268-4684-91b1-9fca91f4e9c1","OpenAI 发布 ChatGPT Images 2.5:画个涂鸦就能出图,生成延迟砍半","openai-chatgpt-images-2-5-sketch","2026-09-09T19:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c480d2d0-9156-4aa7-826f-fba2f252b6b7","蚂蚁开源 LLaDA-Image:6B 参数生成编辑一体,Turbo 版 4 步出图","ant-llada-image-open-generation-editing","2026-09-04T13:09:21+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"5ff06769-2251-40fd-a672-f394a1f68965","十人合影谁是谁:腾讯混元 WithEveryone 给群像生成装上身份锚点","witheveryone-group-image-identity-grounding","2026-08-24T13:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"6082cd23-0eca-40e0-9315-67318dc818ee","NovelAI Diffusion V5 发布:规模翻倍、32 通道 VAE,单次生成整页漫画","novelai-diffusion-v5-release","2026-08-22T13:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00+00:00"]