Image generation's longest-running complaint is not quality but control: you write three lines of prompt, composition is a coin flip, and touching one corner means re-rolling the whole picture. On October 1, Black Forest Labs released FLUX 3 Image, the image component of its multimodal FLUX 3 family, and it makes control a first-class citizen — you draw the boxes, the model paints inside them, and every pixel outside stays put.
Bounding boxes as instructions
FLUX 3 is BFL's multimodal family covering video, audio, images and actions; FLUX 3 Image handles image generation and editing. The core interaction is the bounding box. The canvas is a 0-to-1000 grid on both axes, and each element is declared with a box in [y_min, x_min, y_max, x_max] format plus a text description.
A full layout prompt has two parts: a global caption describing the whole image, followed by a JSON element table where each row carries an id, a bounding box and a description. The caption cites every element by its id — animal_1 and friends — at first mention. This upgrades natural-language generation into structured declaration: position, content and relationships sit explicitly in the table, not guessed from one sentence.
You don't have to draw the boxes yourself. Hand an LLM one line and an aspect ratio, and it plans the caption and element table; every box stays editable afterward. On the model side, a prompt upsampler expands short requests into the dense captions FLUX 3 was trained on, but every box you drew reaches the model verbatim, with the same id and coordinates.
Edit one box, keep the rest
Editing works box by box: re-describe a box, swap its content, or move it — everything untouched stays exactly where it was. BFL calls it pixel-perfect editing: add two divers, replace the duck, print on a shirt, round after round without the image falling apart. TechTimes frames this exact guarantee as the capability gap that has kept AI image editing unreliable in production workflows and LLM-driven pipelines.
Other specs: up to 10 reference images per composition, native full-resolution rendering (an official example shows 5456 × 3072 output), and a commercial weights license that lets companies fine-tune and deploy on their own infrastructure.
An interface built for agents
Seen as a feature update, this is nice; seen as interface design, it is clearly aimed at the agent era. JSON in, JSON out, verifiable coordinates, localized edits — exactly what an LLM pipeline needs when it consumes an image model: an upstream model plans the layout, a downstream model verifies box by box, and every step is checkable and reversible. Plain-prompt interfaces are black boxes by comparison; whether an edit landed correctly needs another vision model to guess.
BFL says it outright on the page: an agent can use the model to create well-composed images with just a text prompt. Pair that with the commercial-weights licensing strategy and the target market is clearly bigger than designer tools — it is companies wiring image generation into automated workflows.
The takeaway: image generation is shifting from a wish-granting well to a plannable canvas. When position can be declared and edits localized, using image models starts to look like layout and engineering rather than gacha. Next time AI cannot edit your picture right, consider whether the problem is that you are still using one sentence — while the interface moved on to tables.
References: official model page https://bfl.ai/models/flux-3-image ; The Decoder https://the-decoder.com/black-forest-labs-launches-flux-3-image-with-multi-step-editing-that-leaves-the-rest-of-your-picture-alone/