InsightsPublished 9 min read

FLUX 3 Image: placing every element exactly where you want it

Most image generators make you describe a picture and hope the pieces land in the right places. Black Forest Labs' new model lets you draw where each element goes, change one part without disturbing the rest, and work at print resolution. Here is how it works and where it fits in a real creative workflow.

By Monolith

A graphic designer at a desktop computer arranging a generated coffee shop scene with selection boxes in a photo editing program
A designer arranging a generated coffee shop scene in a photo editing program.

Picture a coffee roaster in Spokane that needs a fall poster: the headline across the top, the new bag on the left, a barista pouring on the right, a cup in the foreground and a line of small type along the bottom. With most image generators you write that as a paragraph and roll the dice. The bag ends up in the middle, the headline overlaps the barista, and fixing one thing changes three others. That is the problem FLUX 3 Image is built to solve.

Black Forest Labs, the lab behind the FLUX models, launched FLUX 3 Image on October 1, 2026 under the line: control every pixel. It is the image part of FLUX 3, the multimodal model for images, video, audio and robot actions that the company first released in early access in July. 134

Black Forest Labs released a short launch film alongside the model. It plays below in our player; nothing downloads until you press play.

Black Forest Labs' official FLUX 3 Image launch film, released October 1, 2026 (49 seconds).

Bounding boxes: drawing the layout instead of describing it

The headline feature is layout by bounding box. You draw a box for every element that matters and describe what goes inside it, then write one line that ties the scene together. FLUX 3 renders the image with each element inside its box. Whatever aspect ratio you pick, the canvas is a grid from 0 to 1000 on both axes, and each box is written as four numbers: top, left, bottom, right. 1

Behind the scenes, a layout prompt has two parts: a caption for the whole image, and an element table listing each element's id, box and description. Black Forest Labs says the boxes you draw reach the model exactly as written. A prompt upsampler expands your short request into the detailed caption the model was trained on, but it does not move your boxes. 1

A layout prompt, drawn
headline_text_1bag_1barista_1cup_1caption_text_1

Scene promptA warm fall poster for a neighborhood coffee roaster: the new seasonal bag on the left, a barista pouring a latte on the right, a cup in the foreground, headline across the top.

  1. headline_text_1[40, 80, 170, 920]The words "FALL ROAST" in a bold cream sans-serif
  2. bag_1[240, 60, 760, 470]A matte kraft coffee bag with a plain orange label
  3. barista_1[210, 520, 860, 960]A barista pouring a latte, warm window light from the right
  4. cup_1[700, 300, 900, 600]A ceramic cup with latte art on a wooden counter
  5. caption_text_1[920, 80, 980, 920]"Now pouring in Spokane" in a small serif
An illustrative layout prompt written by Monolith for this article, using the box format in Black Forest Labs' documentation: [top, left, bottom, right] on a 0 to 1000 grid. Not a FLUX 3 output. [1][2]

You do not have to draw every box yourself. Give the system one line and an aspect ratio, and a language model plans the layout for you, writing the caption and the element table. Every box stays editable, so you can move the ones you do not like and keep the rest. Black Forest Labs describes the model as designed for agents for exactly this reason: an AI assistant can plan a composition and hand it straight to the image model. 1

The company says boxes suit compositions with many parts in strict relationships: type set around a photograph, collages and panel grids, editorial spreads, or a crowded scene where every face has its place. For a single portrait, a plain text prompt still works, and Black Forest Labs says the model follows prompts well and understands composition on its own. 1

Editing one part without breaking the rest

The second big change is editing. You can re-describe a box, replace what is in it, or move it, and make several of those edits in one request. Everything you did not touch stays where it was. Black Forest Labs' examples include recoloring a surfer's wetsuit and board while the wave and sky stay locked, swapping one character for another, and replacing a drawing on a page while its printed text stays exactly the same. 1

The company calls this pixel-perfect editing, and its launch post puts it as making precise multi-turn edits without changing any other pixel. For brand work that is the difference between a usable tool and a toy. A client who approves a product shot and asks for a different background should get the same product back, not a near-copy with a different label. 13

Over-the-shoulder view of a designer placing outlined layout boxes over a coffee product image in a photo editing program
A designer placing layout boxes over a coffee product image in a photo editing program.

Ten references, one composition

FLUX 3 Image accepts up to 10 reference images and turns them into one composed picture. Each reference gets a token in the order you add it, starting at ref_image_0, and a one-line prompt cites each by its token. The model decides where each item sits and how big it is. Black Forest Labs' own example builds an outfit shot from six separate product photos: a vest, a tee, jeans, a duffel bag, a beanie and sneakers. 12

For retailers and the agencies that serve them, that is a catalog workflow: photograph each product once, then compose lifestyle and lookbook images from the approved shots instead of restaging a shoot. References can be up to 16 megapixels each, so high-resolution product photography goes in without heavy downsizing. 2

Native 4K and the settings that matter

FLUX 3 Image renders natively at up to 4K; Black Forest Labs says it works at full resolution so small details like textures, faces and colors are preserved. One Black Forest Labs sample is 5456 by 3072 pixels, 16.8 megapixels, with hand-lettered characters on a shop sign about 225 pixels tall in the file. That is enough for print and large-format work, not just social posts. 1

SettingOptionsWhat it means for you
Resolution768 square, 1K (default), 2K, 4KDraft small, deliver large
Aspect ratio15 ratios from 21:9 to 9:21, or autoOne request per placement, no cropping
References1 to 10 images, 256 px to 16 MP eachProducts, faces, styles, logos
GroundingOn by defaultSearches the web and images before generating
Safety tolerance0 (strictest) to 4, default 2How permissive moderation is
FLUX 3 Image API settings from Black Forest Labs' documentation, read October 1, 2026. Only the prompt is required. 2

One setting deserves attention. Grounding is switched on by default, which means the model runs a web and image search before it generates. That can help with real places and current objects. It also means outside imagery can influence your result, so for brand work, check the output for anything that resembles someone else's trademark or protected design before it ships. 2

Availability, pricing and weights

FLUX 3 Image is available now through the Black Forest Labs API and the company's Playground, where you can draw boxes by hand. Until October 8, it is 50% off through the API. Black Forest Labs had not published per-image rates for FLUX 3 Image in its launch materials; each API request returns its cost when you submit it, so run a small test before quoting a client. 123

For companies generating images at scale, Black Forest Labs offers FLUX 3 Image under a commercial weights license, so you can fine-tune the model and run it on your own infrastructure. An open-weights version is due in the coming weeks. That matters for agencies with brand-specific styles: a fine-tuned private model can learn a house look that a shared API cannot. 13

FLUX 3 Image is one part of a larger model. The same FLUX 3 model also generates video with native audio, up to 20 seconds in a single generation, which Black Forest Labs released in early access in July. Its official FLUX 3 Video film is below. 4

Black Forest Labs' official FLUX 3 Video film (41 seconds), showing the video side of the same FLUX 3 model.
The FLUX 3 rollout
  1. July 23, 2026FLUX 3 family in early accessReleasedOne multimodal model for images, video with audio, and robot actions.
  2. Oct 1, 2026FLUX 3 Image launchesAvailableAPI and Playground, with commercial weights available to companies.
  3. Until Oct 850% off through the APILimited timeIntroductory discount on API usage.
  4. Coming weeksOpen-weights FLUX 3 ImageAnnouncedA version you can download and run yourself.
Dates and status from Black Forest Labs' FLUX 3 announcement and FLUX 3 Image launch thread. [3][4]

Where it fits in a creative workflow

JobFLUX 3 Image featureWhat to check
Poster or ad with fixed layoutBounding boxes for headline, product and peopleType spelling and brand colors
Product lifestyle imagesUp to 10 product references in one sceneEvery product matches the real item
Client revisionsMulti-edit by box, rest lockedUntouched areas really are unchanged
Print and out-of-homeNative 2K and 4K outputDetail at full size, not just on screen
Campaign at scaleCommercial weights, fine-tuned on house styleLicense terms and hosting cost
Agent-built visualsLayout planned by an AI assistantA person approves before anything publishes
Monolith's suggested uses, mapped to capabilities Black Forest Labs describes. Planning suggestions, not results of Monolith testing. 12

There is no independent benchmark for FLUX 3 Image yet, and Black Forest Labs' July results for the image side were preliminary. So treat this as a capability story, not a quality ranking. The way to judge it is on your own jobs: rebuild one recent layout with boxes, make three rounds of client-style edits, and compare the result and the time spent against the tool you use now. 4

Comparing image tools? Our guides cover Midjourney's editing features and Google's fast Nano Banana option.

The short version

Monolith's take: FLUX 3 Image moves AI imagery from describe-and-hope to design-and-place. Bounding boxes, locked edits and ten-reference composition are the controls agencies have been missing for layout-heavy and brand-sensitive work. Test it on a real layout before the discount ends, and watch the open-weights release if you want a private model trained on your own style.

Questions, answered
What is FLUX 3 Image?
It is the image generation and editing part of Black Forest Labs' FLUX 3 model, launched October 1, 2026. It lets you place elements with bounding boxes, make several edits at once without changing the rest, use up to 10 references and render in up to 4K.
How do bounding boxes work in FLUX 3?
You draw a box for each element and describe it, then add one line for the whole scene. The canvas is a 0 to 1000 grid on both axes, and each box is written as top, left, bottom, right.
How much does FLUX 3 Image cost?
It is available through the Black Forest Labs API at 50% off until October 8, 2026. Per-image rates were not published in the launch materials; each API request returns its cost.
Sources

Read for this feature. The numbers match the markers in the text.

  1. Black Forest Labs: FLUX 3 Imagebfl.ai
  2. Black Forest Labs docs: FLUX 3 Image overviewdocs.bfl.ai
  3. Black Forest Labs on X: Introducing FLUX 3 Image (October 1, 2026)x.com
  4. Black Forest Labs: FLUX 3, Real World Models (July 23, 2026)bfl.ai

Planning a campaign that needs precise layouts, product accuracy and print-ready files? That is what our creative team builds.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.