Bounding-Box Layout Control
Element placement learned during training, not approximated at render time.
A text-to-image foundation model trained from scratch around layout, with bounding-box placement, hex colour conditioning, and native 2K output.
Abre esta página en un navegador de escritorio para empezar a crear.
Ideogram 4.0 approaches type from the opposite direction to most image models, and the difference is architectural rather than stylistic.
It was trained on bounding boxes paired with plain-language descriptions, which means the model learns where each object, text region, and layout element belongs before it renders anything. Ask a photographic model for a gig poster and you generally get a photograph with lettering placed on it. Ask this one and you get a layout — type sized against the space it occupies, imagery arranged to leave room for it, and a hierarchy between headline and supporting copy.
Use Ideogram 4.0 when the deliverable is a piece of design. Produce posters, wordmarks, packaging fronts, ad creative, social graphics, and anything where the words are the design rather than a label attached to one.
Ideogram 4.0 is a text-to-image foundation model trained from scratch, with a feature set shaped around design work rather than around photography.
The model is built around three major strengths:
Output is native 2K, which is print-ready without a separate upscaling pass. Placement can be stated in ordinary language — title at the top, subject centred, logo bottom right — or specified exactly through bounding boxes in a structured JSON brief, which is the reliable route when an element has to sit in a precise position.
Element placement learned during training, not approximated at render time.
Specify exact positions and values rather than describing them.
State a brand colour by value and get that colour.
Print-ready without a separate upscaling step.
Display and supporting type across more than one script.
A foundation model built for this purpose rather than fine-tuned toward it.
Because the training paired boxes with descriptions, position is something the model understands rather than infers. A JSON brief carrying bounding boxes puts an element where you said, which is the difference between generating a poster and designing one.
Hex conditioning removes the round trip where a designer explains that the blue is not the right blue. State the value and it appears, which is what makes campaign-accurate output repeatable across a set.
Headline, subhead, and supporting copy come back at sensibly different weights and sizes, rather than as three strings at similar scale competing for the same attention.
Logo exploration in volume is one of the practical uses. Output is raster rather than vector, so a chosen direction still needs redrawing for production, but finding the direction is the expensive part.
Native 2K means poster and packaging work leaves the model at a usable size instead of needing an upscaling pass that softens the type.
Multilingual typography means a campaign layout can carry a second script without the structure falling apart around it.
Layout work is comparative. Nobody picks a poster direction from one render, and nobody picks a wordmark from fewer than a dozen. Virse lays the whole set out at once. Thirty logo directions sit in a grid you can scan, and the JSON brief that produced them stays beside the results, so amending one value and re-running is a small edit rather than a rebuild.
Generate a spread of layout directions and judge them together, which is the only way this kind of decision actually gets made.
Park the structured brief beside the renders so changing one bounding box or hex value is an edit in place.
Move between Ideogram 4.0 and 30+ other image and video models without leaving the canvas or rewriting the brief.
Send an approved poster into FLUX Kontext for a targeted change, or into a video model as an opening frame.
Gig posters, festival lineups, and exhibition announcements with a working hierarchy.
Dozens of typographic directions to choose between before committing to one.
Product name, descriptor, and volume at plausible relative sizes.
Display ads built to brand colour values from a single structured brief.
Quote cards, announcements, and campaign posts carrying a line of text.
Title-led compositions where type is the primary visual element.
Type every word that should appear, in quotation marks. The model sets what you supply and invents what you do not.
Say which line is the headline and how much larger it should be than the rest.
Describe positions in plain language, or specify bounding boxes and hex values in a JSON brief.
Run a spread and compare, because layout decisions are made by looking at options.
A useful Ideogram 4.0 prompt usually includes four elements:
En lugar de escribir
A cool poster for a jazz concert with the band name and date, nice typography.
Escribe
A concert poster. The headline reads MERIDIAN QUARTET in a tall condensed sans-serif, set in two lines and filling the upper third. Below it, much smaller, 'Thursday 14 November · Union Chapel'. Bottom left corner, smaller again, 'Doors 7pm'. Deep navy background, one warm ochre shape behind the headline, no photography. Wide margins on all four sides.
An exhibition poster in portrait format. Headline reads SLOW WATER in a heavy geometric sans, all caps, two lines, occupying the top quarter. Beneath it in small caps at roughly one fifth the size: "Photographs by Ana Reyes · 3 March – 12 April". Bottom edge, smallest size: "Corran Gallery, Edinburgh". Flat pale grey background with one horizontal deep blue band running behind the headline. No imagery. Generous margin on all sides, all text left-aligned to a common margin.
A wordmark for a coffee roaster called NORTHBOUND. Single line, all caps, a high-contrast serif with pronounced thick-thin transitions and sharp bracketed serifs. Tight letter spacing. Pure black on white, nothing else in frame, no icon, no container shape, no tagline. Centred with equal space above and below.
A front-of-pack layout for a tin of loose leaf tea, shown flat rather than as a product render. Centred and largest: ASSAM SECOND FLUSH. Directly beneath at half the size: "Loose Leaf Black Tea". Bottom edge, small: "100g". Background in #B4593C, one thin cream rule separating the product name from the descriptor. Simple line illustration of a tea leaf above the name, no shading.
| Dimensión | Ideogram 4.0 | Nano Banana Pro |
|---|---|---|
| Approach to type | Composes the layout around it | Renders it inside an image |
| Placement control | Bounding boxes and JSON briefs | Described in the prompt |
| Colour control | Hex value conditioning | Described in the prompt |
| Native output | 2K, print-ready | Up to 4K |
| Best deliverable | Posters, wordmarks, layouts | Photographic images containing words |
| Reference handling | Layout-oriented | Multi-image blending with roles |
The model sets copy you supply and invents copy you do not. Supply all of it, including the small print.
"Half the size of the headline" is actionable. "Smaller" is a guess the model has to make for you.
Specific typeface names are unreliable. Describing the letterform — condensed sans, high-contrast serif, geometric grotesque — is not.
Margin and negative space are half of layout, and a model left to itself will fill them.
Write the words, state the sizes, place the blocks — and the model will set them rather than guess at them.