A strong AI image prompt stacks four to six elements in a consistent order: subject, setting, style, lighting, camera or lens, and aspect ratio. Most disappointing images come from skipping half of those and leaving the model to guess. The more of the stack you fill in — specifically — the closer the result lands to the picture in your head.
The 6-part structure
Think of it as briefing a photographer who can’t see what you’re imagining. Each layer answers a question they’d otherwise have to invent an answer for.
- Subject — who or what, described specifically. Not “a woman” but “a 30-year-old woman with curly red hair in a vintage leather jacket”.
- Setting — where it happens: a rain-slicked Tokyo alley, a sunlit kitchen, an empty studio.
- Style — photorealistic, oil painting, cyberpunk, watercolor, or a named movement or era.
- Lighting — golden hour, soft studio light, harsh neon, moody and dramatic. This sets the whole mood.
- Camera / lens — “85mm, f/1.8”, “wide-angle”, “overhead shot”. Controls depth and framing.
- Aspect ratio — 16:9 for cinematic, 1:1 for social, 9:16 for phone. Say it or the model picks for you.
Why detail beats a three-word prompt
It’s tempting to type “cat astronaut” and hope. But detailed prompts of roughly 30 to 50 words consistently outperform three-word ones — they give you real control over composition, color, and texture instead of a generic average. Detail isn’t clutter here; each specific word is a decision you’re making instead of leaving to the model.
There’s a ceiling, though. Past a point, piling on adjectives starts to fight itself and the model loses the thread. Aim for specific, not maximal — every word should be doing a job.
Midjourney and DALL-E want different things
The same prompt style doesn’t work everywhere. Midjourney responds well to concise, keyword-driven prompts plus its own parameters — “--ar 16:9” for aspect ratio, “--stylize” (0–1000) to control how hard it applies its house aesthetic (low for realism, high for its signature look). DALL-E 3, and the image tools inside ChatGPT and Gemini, do better with full natural-language sentences describing the scene conversationally rather than a stack of comma-separated tokens.
| Tool | What it prefers | Example shape |
|---|---|---|
| Midjourney | Concise keywords + parameters | neon Tokyo alley, rain, cinematic, 85mm --ar 16:9 --stylize 200 |
| DALL-E 3 / ChatGPT | Conversational full sentences | A rainy neon-lit Tokyo alley at night, shot on an 85mm lens, cinematic and moody, wide format. |
| Gemini image | Natural language + clear intent | Create a wide, cinematic photo of a neon Tokyo alley in the rain at night. |
Where this connects to your everyday prompting
If you generate images inside ChatGPT or Gemini — a lot of people do — you’re writing these prompts in the same box where you do everything else. The discipline is identical to text prompting: name the subject, the constraints, and the output shape instead of hoping the model guesses. BeforePrompt runs in those chat boxes and nudges you toward the specific version before you send it.