FLUX 3 Image: placing every element exactly where you want it
Most image generators make you describe a picture and hope the pieces land in the right places. Black Forest Labs' new model lets you draw where each element goes, change one part without disturbing the rest, and work at print resolution. Here is how it works and where it fits in a real creative workflow.
By Monolith

Picture a coffee roaster in Spokane that needs a fall poster: the headline across the top, the new bag on the left, a barista pouring on the right, a cup in the foreground and a line of small type along the bottom. With most image generators you write that as a paragraph and roll the dice. The bag ends up in the middle, the headline overlaps the barista, and fixing one thing changes three others. That is the problem FLUX 3 Image is built to solve.
Black Forest Labs, the lab behind the FLUX models, launched FLUX 3 Image on October 1, 2026 under the line: control every pixel. It is the image part of FLUX 3, the multimodal model for images, video, audio and robot actions that the company first released in early access in July. 134
Black Forest Labs released a short launch film alongside the model. It plays below in our player; nothing downloads until you press play.
Bounding boxes: drawing the layout instead of describing it
The headline feature is layout by bounding box. You draw a box for every element that matters and describe what goes inside it, then write one line that ties the scene together. FLUX 3 renders the image with each element inside its box. Whatever aspect ratio you pick, the canvas is a grid from 0 to 1000 on both axes, and each box is written as four numbers: top, left, bottom, right. 1
Behind the scenes, a layout prompt has two parts: a caption for the whole image, and an element table listing each element's id, box and description. Black Forest Labs says the boxes you draw reach the model exactly as written. A prompt upsampler expands your short request into the detailed caption the model was trained on, but it does not move your boxes. 1
Scene promptA warm fall poster for a neighborhood coffee roaster: the new seasonal bag on the left, a barista pouring a latte on the right, a cup in the foreground, headline across the top.
- headline_text_1[40, 80, 170, 920]The words "FALL ROAST" in a bold cream sans-serif
- bag_1[240, 60, 760, 470]A matte kraft coffee bag with a plain orange label
- barista_1[210, 520, 860, 960]A barista pouring a latte, warm window light from the right
- cup_1[700, 300, 900, 600]A ceramic cup with latte art on a wooden counter
- caption_text_1[920, 80, 980, 920]"Now pouring in Spokane" in a small serif
You do not have to draw every box yourself. Give the system one line and an aspect ratio, and a language model plans the layout for you, writing the caption and the element table. Every box stays editable, so you can move the ones you do not like and keep the rest. Black Forest Labs describes the model as designed for agents for exactly this reason: an AI assistant can plan a composition and hand it straight to the image model. 1
The company says boxes suit compositions with many parts in strict relationships: type set around a photograph, collages and panel grids, editorial spreads, or a crowded scene where every face has its place. For a single portrait, a plain text prompt still works, and Black Forest Labs says the model follows prompts well and understands composition on its own. 1
Editing one part without breaking the rest
The second big change is editing. You can re-describe a box, replace what is in it, or move it, and make several of those edits in one request. Everything you did not touch stays where it was. Black Forest Labs' examples include recoloring a surfer's wetsuit and board while the wave and sky stay locked, swapping one character for another, and replacing a drawing on a page while its printed text stays exactly the same. 1
The company calls this pixel-perfect editing, and its launch post puts it as making precise multi-turn edits without changing any other pixel. For brand work that is the difference between a usable tool and a toy. A client who approves a product shot and asks for a different background should get the same product back, not a near-copy with a different label. 13

Ten references, one composition
FLUX 3 Image accepts up to 10 reference images and turns them into one composed picture. Each reference gets a token in the order you add it, starting at ref_image_0, and a one-line prompt cites each by its token. The model decides where each item sits and how big it is. Black Forest Labs' own example builds an outfit shot from six separate product photos: a vest, a tee, jeans, a duffel bag, a beanie and sneakers. 12
For retailers and the agencies that serve them, that is a catalog workflow: photograph each product once, then compose lifestyle and lookbook images from the approved shots instead of restaging a shoot. References can be up to 16 megapixels each, so high-resolution product photography goes in without heavy downsizing. 2
Native 4K and the settings that matter
FLUX 3 Image renders natively at up to 4K; Black Forest Labs says it works at full resolution so small details like textures, faces and colors are preserved. One Black Forest Labs sample is 5456 by 3072 pixels, 16.8 megapixels, with hand-lettered characters on a shop sign about 225 pixels tall in the file. That is enough for print and large-format work, not just social posts. 1
| Setting | Options | What it means for you |
|---|---|---|
| Resolution | 768 square, 1K (default), 2K, 4K | Draft small, deliver large |
| Aspect ratio | 15 ratios from 21:9 to 9:21, or auto | One request per placement, no cropping |
| References | 1 to 10 images, 256 px to 16 MP each | Products, faces, styles, logos |
| Grounding | On by default | Searches the web and images before generating |
| Safety tolerance | 0 (strictest) to 4, default 2 | How permissive moderation is |
One setting deserves attention. Grounding is switched on by default, which means the model runs a web and image search before it generates. That can help with real places and current objects. It also means outside imagery can influence your result, so for brand work, check the output for anything that resembles someone else's trademark or protected design before it ships. 2
Availability, pricing and weights
FLUX 3 Image is available now through the Black Forest Labs API and the company's Playground, where you can draw boxes by hand. Until October 8, it is 50% off through the API. Black Forest Labs had not published per-image rates for FLUX 3 Image in its launch materials; each API request returns its cost when you submit it, so run a small test before quoting a client. 123
For companies generating images at scale, Black Forest Labs offers FLUX 3 Image under a commercial weights license, so you can fine-tune the model and run it on your own infrastructure. An open-weights version is due in the coming weeks. That matters for agencies with brand-specific styles: a fine-tuned private model can learn a house look that a shared API cannot. 13
FLUX 3 Image is one part of a larger model. The same FLUX 3 model also generates video with native audio, up to 20 seconds in a single generation, which Black Forest Labs released in early access in July. Its official FLUX 3 Video film is below. 4
- July 23, 2026FLUX 3 family in early accessReleasedOne multimodal model for images, video with audio, and robot actions.
- Oct 1, 2026FLUX 3 Image launchesAvailableAPI and Playground, with commercial weights available to companies.
- Until Oct 850% off through the APILimited timeIntroductory discount on API usage.
- Coming weeksOpen-weights FLUX 3 ImageAnnouncedA version you can download and run yourself.
Where it fits in a creative workflow
| Job | FLUX 3 Image feature | What to check |
|---|---|---|
| Poster or ad with fixed layout | Bounding boxes for headline, product and people | Type spelling and brand colors |
| Product lifestyle images | Up to 10 product references in one scene | Every product matches the real item |
| Client revisions | Multi-edit by box, rest locked | Untouched areas really are unchanged |
| Print and out-of-home | Native 2K and 4K output | Detail at full size, not just on screen |
| Campaign at scale | Commercial weights, fine-tuned on house style | License terms and hosting cost |
| Agent-built visuals | Layout planned by an AI assistant | A person approves before anything publishes |
There is no independent benchmark for FLUX 3 Image yet, and Black Forest Labs' July results for the image side were preliminary. So treat this as a capability story, not a quality ranking. The way to judge it is on your own jobs: rebuild one recent layout with boxes, make three rounds of client-style edits, and compare the result and the time spent against the tool you use now. 4
Comparing image tools? Our guides cover Midjourney's editing features and Google's fast Nano Banana option.
Monolith's take: FLUX 3 Image moves AI imagery from describe-and-hope to design-and-place. Bounding boxes, locked edits and ten-reference composition are the controls agencies have been missing for layout-heavy and brand-sensitive work. Test it on a real layout before the discount ends, and watch the open-weights release if you want a private model trained on your own style.
What is FLUX 3 Image?
How do bounding boxes work in FLUX 3?
How much does FLUX 3 Image cost?
Read for this feature. The numbers match the markers in the text.
Planning a campaign that needs precise layouts, product accuracy and print-ready files? That is what our creative team builds.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.