All posts

How to write AI video prompts for Sora, Veo, and Runway

August 4, 2026 · 7 min read · By Khalid

TL;DR

  • Describe visible action, camera movement, subject continuity, duration, aspect ratio, style, and audio.
  • Default to slow, smooth camera moves — fast motion still warps subjects in 2026 models.
  • Sora, Veo, and Runway each reward a slightly different prompt order.
  • One clear action beats five vague ones; the model can only track so much per shot.

A good AI video prompt describes visible action, camera movement, subject continuity, duration, aspect ratio, style, and any audio you need. Video adds two things a still image doesn’t: motion and time. If you don’t direct them, the model invents both — usually not the way you hoped. Treat the prompt like a shot list you’re handing a camera crew.

What to put in the prompt

Name the action first — what actually moves and happens on screen. Then the camera: is it a slow push-in, a drone pull-back, a static locked-off shot? Then keep the subject consistent (“the same woman from the first shot”), state a duration, an aspect ratio, and the visual style. Add audio only if the tool supports it.

A prompt like “a drone shot slowly pulling back from a neon-lit Tokyo street at 3 a.m., light rain on the lens” works because it names the shot, the motion speed, the subject, the setting, and the mood — a crew could shoot it from that sentence.

Keep the camera slow — this is the big one

Here’s the mistake almost everyone makes: asking for fast, dramatic camera moves. AI video models still degrade at fast camera speeds — the subject loses detail or warps. So “slow” and “smooth” are the right defaults for most product, narrative, and lifestyle shots. You get a cleaner, more usable clip from a gentle push-in than from a whip-pan that turns a face into soup.

Each model wants a slightly different order

The big 2026 models — Sora 2, Google Veo 3, Runway Gen-4, Kling — respond best to slightly different structures. It’s worth matching the order to the tool.

ModelPrompt order it likesNotable strength
Sora 2subject → action → camera → environment → lighting → mood → styleCoherent, natural motion
Veo 3subject+action → camera → environment+lighting → style+mood → audioBuilt-in audio layer
Runway Gen-4action-first, then granular camera controlsFine camera control: pan, focal length, dolly speed, rack focus
Prompt ordering by model

One action per shot

Don’t cram a whole scene into one prompt. The model can only track so much continuity per clip, and asking for “she walks in, sits down, opens a laptop, and the camera spins around her” usually produces a mess. Storyboard it as separate shots and generate them one at a time. It’s slower, but it’s the difference between usable footage and a warped blur.

The through-line with every other kind of prompting: be specific about what you want and what defines a good result, before you generate. If you draft these prompts inside ChatGPT or Gemini before pasting them into a video tool, BeforePrompt is right there to catch the vague version first.

Frequently asked

What should an AI video prompt include?

Visible action, camera movement, subject continuity, duration, aspect ratio, style, and any needed audio. Video adds motion and time to an image prompt — if you don’t direct them explicitly, the model invents both.

Why do my AI videos look warped or blurry?

Usually fast camera movement. AI video models in 2026 still degrade quality at fast camera speeds — subjects lose detail or warp. Default to slow, smooth moves like a gentle push-in or pull-back, and generate one clear action per shot rather than a whole scene at once.

Do Sora, Veo, and Runway need different prompts?

The order helps. Sora 2 likes subject, action, camera, environment, lighting, mood, style. Veo 3 works in layers ending with audio. Runway Gen-4 is action-first with granular camera controls like pan, focal length, and dolly speed.

Catch the gap before you send

BeforePrompt flags the missing piece while the prompt is still yours to change.