Image Generation Prompting Tips: 10 Fixes I Tested Across 4 Models

Ten image generation prompting tips, each shown as a before and after pair generated on Nano Banana 2 Lite, Seedream 5 Pro, FLUX.2 Pro and Ideogram 4.

A designer's night workspace lit by a single warm pendant lamp, with the title Image Prompting

Most bad AI images are not a model problem. I ran the same set of tests across four image models on a single afternoon, and in almost every case the difference between a throwaway image and one I would actually ship came from six or seven extra words in the prompt, not from switching models.

So this is not a listicle of prompt "magic words". Every tip below is a before and after pair I generated on Segmind while writing this post, on Nano Banana 2 Lite, Seedream 5 Pro, FLUX.2 Pro and Ideogram 4. The lazy prompt is the one most people actually type. The tested prompt is what I would send in production. Two of the tips exist because a test failed and taught me something I did not expect.

How I ran these tests

Every image in this post is a fresh generation from this run, made through the Segmind API with default parameters unless the tip is specifically about a parameter. I picked four models on purpose: a fast cheap general model (Nano Banana 2 Lite, $0.042 per image), a photographic flagship (Seedream 5 Pro, $0.05625 at 1K), a materials and detail specialist (FLUX.2 Pro, $0.0375 for the first megapixel) and a typography specialist (Ideogram 4, $0.10 per megapixel at QUALITY). Total spend for the whole test set was under two dollars, which is the real reason prompt iteration beats model shopping: you can afford to run the same idea eight ways.

1. Name the shot, not the vibe

"A beautiful photo of X" gives the model no camera to stand behind, so it defaults to a wide, evenly lit, stock-catalogue view of the scene. Tell it where the camera is, how far away, what is in focus and what falls off.

Lazy prompt a beautiful photo of a coffee shop

Tested prompt Straight-on 35mm shot of a corner espresso bar at 7am, framed at chest height, shallow depth of field at f/2 with focus on the portafilter in the barista's hands, the back counter falling into soft blur, warm window light entering from camera left

Parameters model: nano-banana-2-lite  |  aspect_ratio: 3:2  |  thinking_level: high  |  seed: 42

Lazy prompt

Nano Banana 2 Lite output, generic wide cafe interior from a vague prompt

Tested prompt

Nano Banana 2 Lite output, tight 35mm barista shot with shallow depth of field

Same model, same settings. The only change is that the second prompt places a camera in the room.

The first image is a room. The second is a photograph of a moment inside a room, with a subject, a focal plane and a light direction. Four pieces of vocabulary do almost all the work here: focal length, camera height, aperture and where the light comes from. If you only remember one tip from this post, use this one.

2. Put the subject first, put the style last

A lot of prompts still start with a pile of quality tokens inherited from 2023 Stable Diffusion habits: "cinematic, moody, ultra detailed, 8k, hyperrealistic, award winning". Modern models read your prompt as a sentence, so whatever comes first anchors the composition. Lead with the quality tokens and you are telling the model that mood is the subject.

Lazy prompt cinematic, moody, ultra detailed, 8k, hyperrealistic, award winning, trending, a woman in a market

Tested prompt A woman in her fifties weighing tomatoes on a hanging brass scale at a covered vegetable market stall, both hands steadying the pan, wooden crates of produce stacked behind her. Style: documentary photography, muted greens and deep reds, natural overhead daylight through a corrugated roof.

Parameters model: seedream-5-pro  |  size: 1K  |  aspect_ratio: 3:2

Style first

Seedream 5 Pro output, generic moody market scene

Subject first, style last

Seedream 5 Pro output, documentary style portrait of a woman weighing tomatoes

Seedream 5 Pro. The lazy version renders the mood faithfully and forgets to give the woman anything to do.

In the first image the woman is small, turned away and doing nothing in particular, because "moody" got top billing. In the second she has an action, a prop and a place in the frame, and the style instruction still lands because it is attached at the end as a separate sentence. My rule: subject and action, then setting, then camera, then style. Quality tokens like "8k" and "award winning" contribute almost nothing on 2026 models and cost you prompt real estate.

3. Describe the light source, not the mood

"Moody" is an outcome. Models cannot render an outcome, but they can render a lamp. Name the fixture, its position, its colour temperature and what happens to everything it does not reach.

Lazy prompt moody portrait of a chef in a kitchen

Tested prompt Portrait of a chef standing at the pass in a dark prep kitchen, lit only by a single overhead heat lamp directly above him, warm 2700K pool of light falling on his forehead, shoulders and the steel counter, the rest of the kitchen dropping to near black, no fill light

Parameters model: nano-banana-2-lite  |  aspect_ratio: 3:2  |  thinking_level: high

"Moody"

Nano Banana 2 Lite output, evenly lit kitchen portrait

One named light source

Nano Banana 2 Lite output, chef lit by a single overhead heat lamp against a black background

Naming the fixture and killing the fill light is what produces contrast. The word moody produces even ambient light with a warm grade.

Three phrases carry this: "lit only by", a colour temperature in kelvin, and "no fill light". The same pattern works in reverse for bright commercial work: "large softbox at 45 degrees camera left, white bounce card opposite, no shadows under the product".

4. Quote your text exactly, and keep it short

If you describe the text you want instead of writing it, the model will invent copy, and invented copy is where garbled lettering shows up. Put every string you actually want inside quotes, say where it sits in the layout, and keep each string under about six words.

Lazy prompt a poster for a coffee brand with the brand name and a tagline about morning energy

Tested prompt Minimal print poster on a cream background, one espresso cup centred low in the frame, headline text reading "MORNING FOLD" in bold condensed sans across the top, small caption reading "ROASTED IN GOA" at the bottom edge, generous empty space between them

Parameters model: ideogram-4  |  rendering_speed: QUALITY (default)  |  output_format: png

Described text

Ideogram 4 output, invented coffee brand poster with a garbled line of text

Quoted text

Ideogram 4 output, minimal poster reading MORNING FOLD and ROASTED IN GOA

Ideogram 4. Left: the model invents a brand, a tagline and one line of broken lettering under the logotype. Right: both quoted strings render exactly.

Look closely at the left poster: the invented brand mark reads cleanly, but the small line under it dissolves into letter shapes that are not words. That is the failure mode you are avoiding. On the right, "MORNING FOLD" and "ROASTED IN GOA" are both exact, because they were quoted and short. Long paragraphs of body copy still break on every model I tested. If you need real body copy, generate the image with empty space and set the type in your design tool.

5. Name materials, not adjectives

"Premium", "luxury" and "high end" are price signals, not visual instructions. Materials and finishes are visual instructions.

Lazy prompt a luxury watch product shot, premium, high end, expensive looking

Tested prompt Product shot of a wristwatch: brushed titanium case, sapphire crystal carrying a faint blue anti-reflective cast, matte vulcanised rubber strap, resting on honed black basalt, a single softbox reflection running the length of the case, dust-free surface

Parameters model: flux-2-pro  |  width: 1216  |  height: 832  |  seed: 42

Adjectives

FLUX.2 Pro output, generic silver luxury watch product shot

Materials

FLUX.2 Pro output, titanium watch with rubber strap on black basalt

FLUX.2 Pro. Adjectives give you the generic idea of an expensive watch. Materials give you a specific object you could actually photograph.

FLUX.2 Pro rewards this more than any other model I tested. Brushed versus polished, matte versus lacquered, honed versus glossy: each of those pairs changes how the surface handles the light, and the model renders the difference. The same trick works for food (crumb structure, glaze, sear), fabric (slub, ribbed, boiled wool) and architecture (board formed concrete, weathered corten).

6. Set the aspect ratio at generation time, not in the crop tool

Cropping a 1:1 image to 9:16 throws away two thirds of the pixels and usually cuts through the subject. Every model here recomposes the scene when you change the aspect ratio, so ask for the shape you are going to ship.

Lazy prompt A pair of trail running shoes on wet slate rock beside a mountain stream, early morning light, product photography for an outdoor brand (generated at 1:1)

Tested prompt Same prompt, aspect_ratio switched to 9:16 for a story or reel placement

Parameters model: nano-banana-2-lite  |  aspect_ratio: 1:1 vs 9:16  |  seed: 42

1:1

Nano Banana 2 Lite output, square composition of trail shoes by a stream

9:16

Nano Banana 2 Lite output, vertical composition of the same scene

Not a crop. The vertical version moves the subject, changes the amount of stream in frame and keeps the mountains, which a crop of the square image could not do.

If you need one image in several placements, generate it once per placement instead of cropping. At $0.042 per image on Nano Banana 2 Lite, three placements cost about 13 cents, which is cheaper than the ten minutes you would spend fixing crops.

7. Do not expect the seed to lock your composition

This is the tip that came out of a failed test. Old diffusion habits say: fix the seed, change one word, and everything else stays put. I ran exactly that experiment, seed 777 on both, changing only the jacket colour.

Lazy prompt A cyclist standing beside a touring bike on a coastal road at dawn, wearing a red windbreaker, wide shot, sea on the right

Tested prompt A cyclist standing beside a touring bike on a coastal road at dawn, wearing a yellow windbreaker, wide shot, sea on the right

Parameters model: nano-banana-2-lite  |  seed: 777 on both  |  aspect_ratio: 3:2  |  one word changed

seed 777, red

Nano Banana 2 Lite output, cyclist in a red jacket on a coastal road

seed 777, yellow

Nano Banana 2 Lite output, cyclist in a yellow jacket on a different coastal road

Same seed, one word different. The road, the rider, the pose and the light all changed.

Both images are good. Neither is a controlled variation of the other. On the Gemini family models the seed is a nudge toward reproducibility, not the deterministic lock it is on classic diffusion samplers, and any change to the prompt can move the whole frame. If you need a true single-variable change, do it as an edit: pass your first image back in through image_urls and ask for the jacket colour to change while everything else stays the same. Plan your workflow around edits, not seeds.

8. Never ask for "an advertisement"

Ask for an ad and you get an ad, complete with invented headline, invented logo, invented body copy and a call to action button. It looks impressive in a demo and it is unusable, because your designer now has to paint out someone else's typography.

Lazy prompt an advertisement image of running shoes

Tested prompt Advertising still life: one running shoe placed in the lower right third of the frame on a seamless mid-grey backdrop, the upper left two thirds left as clean empty wall for headline copy, soft directional light from the top right, small hard shadow under the shoe

Parameters model: seedream-5-pro  |  size: 1K  |  aspect_ratio: 3:2

"An advertisement"

Seedream 5 Pro output, complete fake advertisement layout with invented brand and headline

Ask for the negative space

Seedream 5 Pro output, product still life with empty negative space for headline copy

Seedream 5 Pro. The right hand frame is the one a designer can actually use, because the copy area is empty and lit evenly.

Describe the plate, not the poster. Say which third the product sits in, say what the rest of the frame is, and set the type yourself. One more thing worth knowing: when I asked for a plain running shoe, the model put a recognisable sportswear logo on it. If the output is going anywhere near a paying client, add "unbranded, no logos, no visible branding" to product prompts and check the result before it ships.

9. Count out loud

Quantity words like "a few", "several" and "some" are read as "a pile of". If the number matters, state it and describe the arrangement.

Lazy prompt a few croissants on a tray

Tested prompt Exactly three croissants in a single row on a steel bakery tray, evenly spaced with a hand's width between each, shot top down, nothing else on the tray

Parameters model: nano-banana-2-lite  |  aspect_ratio: 3:2  |  thinking_level: high

"A few"

Nano Banana 2 Lite output, a pile of croissants on a wooden tray in a bakery

"Exactly three", arrangement described

Nano Banana 2 Lite output, exactly three croissants in a row on a steel tray

Counting works when you also describe the arrangement. Exactly three, in a row, evenly spaced, nothing else in frame.

Accuracy still degrades past about six objects on every model I have tested, so if you need nine identical items in a grid, generate three and composite, or use a workflow that lays them out for you. The same discipline applies to people: "three people at a table" behaves, "a group of people" does not.

10. Draft on the cheap model, finish on the expensive one

This is a workflow tip rather than a wording tip, and it is the one that saves the most money. Prompt iteration is a search problem, so run the search on the cheapest model that understands language well, then re-run the winning prompt on the model whose output quality you actually need.

Here is what the four models in this post cost per image on Segmind today:

Model Price per image Best for
Nano Banana 2 Lite$0.042 flatPrompt iteration, thumbnails, bulk drafts
FLUX.2 Pro$0.0375 first megapixel, $0.01875 each additionalMaterials, surfaces, product detail
Seedream 5 Pro$0.05625 at 1K, $0.1125 at 2KPhotographic scenes, people, campaign stills
Ideogram 4$0.03 per MP TURBO, $0.06 BALANCED, $0.10 QUALITYPosters, packaging, anything with type
Nano Banana 2$0.08 at 1K, $0.12 at 2KInstruction following, composition heavy briefs

Segmind list prices at the time of writing. Charges are per successful generation.

Twenty drafts on Nano Banana 2 Lite cost 84 cents. Twenty drafts at 2K on a flagship cost about four times that, for images you are going to throw away anyway. Iterate cheap, finish expensive.

Where this approach falls short

Longer prompts are not automatically better prompts. Past roughly 120 words I start seeing models drop the third or fourth instruction, and the fix is to move the dropped detail into an edit pass rather than to shout louder in the original prompt. Text longer than a few words still breaks everywhere. Exact brand colours are unreliable, so composite your brand elements rather than prompting for a hex value. And models go down: Nano Banana 2 was returning errors from its upstream provider during this test run, which is why the Lite model does most of the work in this post. If image generation sits in a production path, write the fallback model into your code before you need it.

FAQ

What are the most important image generation prompting tips for 2026?

Name the camera and framing, describe the light source instead of the mood, use materials rather than adjectives, quote any text you want rendered, state exact counts, and set the aspect ratio at generation time. Those six changes fixed more images in my tests than any model switch.

Do quality tokens like "8k, ultra detailed, award winning" still work?

Not meaningfully. On the four models I tested, quality tokens mostly displaced useful description. Modern image models read prompts as language, so the words that pay off are the ones describing the subject, the light and the camera.

Why does the same seed give me a different image?

On Gemini family models such as Nano Banana 2 and Nano Banana 2 Lite, the seed nudges reproducibility but does not lock composition, and any prompt change can move the frame. For controlled single variable changes, pass the original image back in as a reference and request an edit instead.

How do I get accurate text in an AI generated image?

Put the exact string in quotes, keep it under about six words per line, say where each line sits in the layout, and use a typography focused model such as Ideogram 4. For paragraphs of copy, generate empty space and set the type in your design tool.

How many words should an image prompt be?

Between about 30 and 90 words for most work. That is enough for subject, action, setting, light and camera without hitting the point where models start dropping instructions.

What is the cheapest way to iterate on prompts?

Run your search on a low cost model, then re-run the winning prompt on your quality model. Nano Banana 2 Lite is $0.042 per image on Segmind, so twenty iterations cost under a dollar.

Try these on your own prompts

Take one prompt you are unhappy with and apply three of these tips: put the subject first, name the light source, and add the camera. Run it on Nano Banana 2 Lite for four cents, and once the wording is right, re-run it on Seedream 5 Pro or FLUX.2 Pro. All the models in this post run behind one API key on Segmind, so switching between them is a one line change.