All posts

How to write AI image prompts that look how you imagined

August 4, 2026 · 7 min read · By Khalid

TL;DR

  • Stack the prompt: subject, setting, style, lighting, camera/lens, aspect ratio.
  • Be specific about the subject — “a person” gives you a stranger; describe who.
  • Detailed prompts (30–50 words) consistently beat three-word ones for control.
  • Midjourney likes concise keywords + parameters; DALL-E and Gemini like natural sentences.

A strong AI image prompt stacks four to six elements in a consistent order: subject, setting, style, lighting, camera or lens, and aspect ratio. Most disappointing images come from skipping half of those and leaving the model to guess. The more of the stack you fill in — specifically — the closer the result lands to the picture in your head.

The 6-part structure

Think of it as briefing a photographer who can’t see what you’re imagining. Each layer answers a question they’d otherwise have to invent an answer for.

  • Subject — who or what, described specifically. Not “a woman” but “a 30-year-old woman with curly red hair in a vintage leather jacket”.
  • Setting — where it happens: a rain-slicked Tokyo alley, a sunlit kitchen, an empty studio.
  • Style — photorealistic, oil painting, cyberpunk, watercolor, or a named movement or era.
  • Lighting — golden hour, soft studio light, harsh neon, moody and dramatic. This sets the whole mood.
  • Camera / lens — “85mm, f/1.8”, “wide-angle”, “overhead shot”. Controls depth and framing.
  • Aspect ratio — 16:9 for cinematic, 1:1 for social, 9:16 for phone. Say it or the model picks for you.

Why detail beats a three-word prompt

It’s tempting to type “cat astronaut” and hope. But detailed prompts of roughly 30 to 50 words consistently outperform three-word ones — they give you real control over composition, color, and texture instead of a generic average. Detail isn’t clutter here; each specific word is a decision you’re making instead of leaving to the model.

There’s a ceiling, though. Past a point, piling on adjectives starts to fight itself and the model loses the thread. Aim for specific, not maximal — every word should be doing a job.

Midjourney and DALL-E want different things

The same prompt style doesn’t work everywhere. Midjourney responds well to concise, keyword-driven prompts plus its own parameters — “--ar 16:9” for aspect ratio, “--stylize” (0–1000) to control how hard it applies its house aesthetic (low for realism, high for its signature look). DALL-E 3, and the image tools inside ChatGPT and Gemini, do better with full natural-language sentences describing the scene conversationally rather than a stack of comma-separated tokens.

ToolWhat it prefersExample shape
MidjourneyConcise keywords + parametersneon Tokyo alley, rain, cinematic, 85mm --ar 16:9 --stylize 200
DALL-E 3 / ChatGPTConversational full sentencesA rainy neon-lit Tokyo alley at night, shot on an 85mm lens, cinematic and moody, wide format.
Gemini imageNatural language + clear intentCreate a wide, cinematic photo of a neon Tokyo alley in the rain at night.
Same idea, two prompt styles

Where this connects to your everyday prompting

If you generate images inside ChatGPT or Gemini — a lot of people do — you’re writing these prompts in the same box where you do everything else. The discipline is identical to text prompting: name the subject, the constraints, and the output shape instead of hoping the model guesses. BeforePrompt runs in those chat boxes and nudges you toward the specific version before you send it.

Frequently asked

What is the best structure for an AI image prompt?

Stack four to six elements in order: subject, setting, style, lighting, camera or lens, and aspect ratio. Fill each one in specifically. Skipping layers is the most common reason a generated image doesn’t match what you pictured.

How long should an image generation prompt be?

Detailed prompts of roughly 30 to 50 words consistently outperform very short ones — they give you control over composition, color, and texture. But don’t pile on endless adjectives; past a point they fight each other. Aim for specific, not maximal.

Do Midjourney and DALL-E need different prompts?

Yes. Midjourney prefers concise, keyword-driven prompts plus its own parameters like --ar and --stylize. DALL-E 3 and the image tools in ChatGPT and Gemini do better with full natural-language sentences describing the scene conversationally.

Catch the gap before you send

BeforePrompt flags the missing piece while the prompt is still yours to change.