A good AI video prompt describes visible action, camera movement, subject continuity, duration, aspect ratio, style, and any audio you need. Video adds two things a still image doesn’t: motion and time. If you don’t direct them, the model invents both — usually not the way you hoped. Treat the prompt like a shot list you’re handing a camera crew.
What to put in the prompt
Name the action first — what actually moves and happens on screen. Then the camera: is it a slow push-in, a drone pull-back, a static locked-off shot? Then keep the subject consistent (“the same woman from the first shot”), state a duration, an aspect ratio, and the visual style. Add audio only if the tool supports it.
A prompt like “a drone shot slowly pulling back from a neon-lit Tokyo street at 3 a.m., light rain on the lens” works because it names the shot, the motion speed, the subject, the setting, and the mood — a crew could shoot it from that sentence.
Keep the camera slow — this is the big one
Here’s the mistake almost everyone makes: asking for fast, dramatic camera moves. AI video models still degrade at fast camera speeds — the subject loses detail or warps. So “slow” and “smooth” are the right defaults for most product, narrative, and lifestyle shots. You get a cleaner, more usable clip from a gentle push-in than from a whip-pan that turns a face into soup.
Each model wants a slightly different order
The big 2026 models — Sora 2, Google Veo 3, Runway Gen-4, Kling — respond best to slightly different structures. It’s worth matching the order to the tool.
| Model | Prompt order it likes | Notable strength |
|---|---|---|
| Sora 2 | subject → action → camera → environment → lighting → mood → style | Coherent, natural motion |
| Veo 3 | subject+action → camera → environment+lighting → style+mood → audio | Built-in audio layer |
| Runway Gen-4 | action-first, then granular camera controls | Fine camera control: pan, focal length, dolly speed, rack focus |
One action per shot
Don’t cram a whole scene into one prompt. The model can only track so much continuity per clip, and asking for “she walks in, sits down, opens a laptop, and the camera spins around her” usually produces a mess. Storyboard it as separate shots and generate them one at a time. It’s slower, but it’s the difference between usable footage and a warped blur.
The through-line with every other kind of prompting: be specific about what you want and what defines a good result, before you generate. If you draft these prompts inside ChatGPT or Gemini before pasting them into a video tool, BeforePrompt is right there to catch the vague version first.