Video Prompt Fundamentals — Shot Language, Motion, and Duration
A video prompt is not an image prompt with 'video' tacked on. It needs a subject, a camera, a style, and a duration — in that order — or the model decides for you.
Image prompts and video prompts look similar on the surface — both describe a scene, both reward specificity. But a video has a dimension an image doesn't: time. A camera has a starting position and a way it moves through the shot. A subject has a trajectory. Get the words right but skip the motion, and the output looks flat — not because the platform failed, but because you never told it to do anything else.
The Four Layers
This track works with four layers. Stack all four into every prompt, in order, and you'll rarely land on the model's boring default.
Scene & Subject — who or what is in frame, and where. Be as concrete as you'd be describing a photograph to someone who can't see it: not "a person outdoors" but "a lone hiker crossing a ridge at dawn, fog below in the valley."
Camera & Motion — how the camera behaves, and how the subject moves. This is the layer most learners skip, and it's the one that separates a video from an animated photograph. Learn a handful of terms and reuse them constantly:
- Static — the camera doesn't move at all
- Pan — the camera rotates left or right from a fixed point
- Dolly in / dolly out — the camera physically moves toward or away from the subject
- Tracking shot — the camera follows a moving subject, keeping pace
- Push in — a slow, subtle dolly used for emphasis, often on a face
Style & Atmosphere — lighting, color grade, mood, and any film reference you want to invoke. "Golden-hour light, muted cool tones, quiet and vast" tells the model far more than "cinematic."
Technical Constraints — duration and aspect ratio, stated explicitly. Leaving these out doesn't mean the model asks you — it picks a default, usually a short square-ish clip that may not fit what you're building.
A Worked Sequence
Here's what stacking all four layers looks like across a short, multi-shot sequence — the kind of thing you'll build in the next lesson.
Notice that duration is a per-shot decision, not a per-sequence one. A "15-second video" in this framework is three shots — 4, 6, and 5 seconds — generated independently and cut together, not one long continuous call. That constraint isn't a limitation to work around; it's the actual unit you're writing for. Every prompt you write specifies the duration of that one shot.
Shot Language You'll Actually Use
You don't need a film-school vocabulary — a handful of terms cover most of what a beginner needs:
- Wide / establishing shot — shows the full scene and where things are relative to each other; usually your opening shot
- Medium shot — frames a subject from roughly the waist up; the workhorse for most narrative content
- Close-up — frames a face or a small detail; used for emphasis, not for every shot
- Tracking shot — as above, following movement; good for a sense of momentum
- Push in — as above, a slow approach; good for a moment of realization or emotional weight
Pair one shot type per generation. Asking for "a wide shot that becomes a close-up" in a single prompt is asking for a camera behavior most platforms handle poorly in one continuous take — that's a two-shot sequence, not one prompt.
Duration Is a Constraint, Not an Afterthought
Every platform has a practical ceiling on a single generation — typically a handful of seconds, sometimes reaching into the twenties on platforms built for longer narrative shots. Two consequences follow directly:
First, always state duration explicitly. An unstated duration gets you the platform's default, which is rarely the length you actually wanted.
Second, plan longer content as a sequence of shots from the start, not as one ambitious prompt that will disappoint you. The shot-by-shot habit you build in this lesson is the same habit a full production pipeline runs on — you're just doing it manually here, one clip at a time.
Lesson Drill
Take one idea for a video and write it out across all four layers, on paper or in a notes app, before you touch any generation tool. Read it back. If any layer reads vague — "cinematic," "nice lighting," "cool camera move" — that's the exact spot the model will fill in for you. Fix it before you spend a generation on it.