AI VideoPromptingCinematographyMedia Generation

Camera control with Wan: the vocabulary that actually moves a camera

Flat generated footage is almost never a model problem. It is a direction problem. The prompt described a picture and left the camera to invent its own behaviour, so the camera invented drift. Here is the vocabulary that replaces drift with intent, the six moves worth learning, and the phrasing that makes each one land.

2026-08-12·10 min read

“Cinematic” is a review, not an instruction

The most common first attempt looks like this: cinematic shot of a lighthouse at dusk, dramatic, film look, 4k. What comes back is a lighthouse at dusk with a camera wandering gently around it, because nothing in that sentence says where the camera starts, where it ends, or how fast it travels between the two.

Every word in that prompt is about how you want to feel about the result. None of it is about what the camera does. Words like cinematic, epic, dramatic and beautiful are the note you give after a screening. They are not the note you give the operator before the take. Cut them entirely and spend the same space on geometry and pace, and the output changes more than any parameter you can tune.

A shot and a move are two different sentences

There are two things to describe, and they answer different questions. The shot is what is in frame: subject, framing, light, background, wardrobe. The move is what the camera does across the duration: direction, distance, speed.

Most prompts describe the shot in loving detail and say nothing about the move, then treat whatever motion appears as the model's fault. Motion always appears. Video output has a time axis and something has to fill it. When nothing constrains that axis you get the default: a slow lateral float, a breathing zoom, a handheld wobble nobody asked for. Drift is not a glitch. Drift is what motion looks like when it is unconstrained.

So write them as separate clauses, in order, and stop mixing them. Shot first, move second, constraints third.

  1. 01
    State the shot
    Subject, framing, light, location. One sentence. No camera language in it yet.
  2. 02
    Name one move
    Grip language, not adjectives: push in, orbit left, tilt up. One move per shot.
  3. 03
    Fix the endpoints
    Where the framing starts and where it ends. Wide to medium close-up, floor to skyline.
  4. 04
    Set the pace
    Steady walking pace, slow, one continuous take. Pace is what separates a push from a lunge.
  5. 05
    Rule out the rest
    No handheld shake, no zoom, no cut, no second subject. Short, camera-focused negatives.
The order matters more than the wording. Every step narrows the space of motions the shot could contain, so the last step has very little left to rule out.

The six moves worth knowing

Almost everything you will cut into a sequence is one of six moves, or a locked-off frame that deliberately refuses to be one. Learn the phrasing for these and you can direct the large majority of shots without reaching for anything exotic.

MoveWhat to writeWhat goes wrong without it
Push inA slow dolly push toward [subject], starting wide and ending at a medium close-up, steady walking pace.A zoom instead of a dolly, so the background flattens rather than compressing, and the shot overshoots into a face-filling close-up.
Pull outThe camera pulls back from [subject] to reveal [context], ending on a wide that holds the whole [location] in frame.The pull keeps going past the reveal and ends on an empty wide with the subject lost in it.
OrbitOrbit the camera 90 degrees to the left around [subject], keeping it centred and the same size throughout.A partial swing that stalls halfway, or an orbit that drifts closer as it travels and turns into a spiral.
PanThe camera pans right from [starting element] to [ending element] on a fixed tripod head, no lateral travel.The camera translates sideways instead of rotating, which reads as a track and breaks the parallax you wanted.
TiltThe camera tilts up from [ground detail] to [upper element], pivoting in place, ending with [element] centred.A crane move: the whole camera rises, so the low-angle geometry that made the tilt worth doing disappears.
Locked-offStatic locked-off camera on a tripod. Nothing moves except [one small motion]. No camera drift.The default float. A shot that was supposed to be the calm anchor in the cut breathes through the whole take.

Endpoints are what stop the overshoot

“Push in on the subject” names a direction and no destination. A direction with no destination runs until the clip ends, which is how you get shots that start nicely composed and finish on a nostril.

Naming the end framing turns an open-ended direction into a bounded one. Starting wide and ending at a medium close-up gives the move a finish line. So does ending with the horizon on the upper third, or ending back where it started for a full orbit. The vocabulary here is the standard framing ladder: wide, medium wide, medium, medium close-up, close-up, extreme close-up. Use two of those rungs per shot, one for the start and one for the end, and the distance travelled is decided before the render begins rather than during it.

The same trick fixes pace. A move with a named start and end has an implied speed, and adding steady or slow, no acceleration pins it down. Most shots that feel cheap are moving too fast for the distance they cover.

Negative constraints do more work than the description

Naming one move does not forbid the others. A push-in prompt can still come back with a push-in plus a handheld tremor plus a slow roll, because none of the extras were excluded. Four short negatives handle nearly all of it:

  • No handheld shake. Removes the operator-breathing texture that gets added to anything vaguely documentary.
  • No camera drift. The one that makes locked-off shots actually lock off.
  • No zoom. Keeps a dolly a dolly, so the background compresses the way a physical push does.
  • No cut, no scene change. Stops a single shot turning into a two-shot mini-edit halfway through the duration you paid for.

Keep negatives about camera behaviour, short, and few. A long list of things you do not want is another way of describing a picture, and it competes with the move for attention. Four is usually enough. If you find yourself writing eight, the shot is probably trying to do two things and wants splitting.

Let the image hold the art direction

Text-to-video asks one prompt to carry the look and the motion at once, and the two compete. Every extra clause about wardrobe, palette and lens is a clause that is not about the camera.

Image-to-video splits the job. A still you already approved carries the composition, the palette, the lighting direction and the character, so the prompt is free to be almost entirely about motion. Orbit the camera 90 degrees to the left around the subject, same size throughout, parallax reveals the doorway behind. No zoom, no cut. That is a complete instruction when the first frame has already settled everything else, and it is a much shorter thing to iterate on when the move is not quite right.

On RenderHour, Wan 2.7 ships text-to-video, image-to-video and audio-to-video variants, so the same family covers a still-driven shot and a text-driven one without changing your vocabulary between them. LTX-2.3 adds a video-to-video path when the input is already footage, and HunyuanVideo 2.0 covers text and image entry points as well. The whole open-weight catalogue is hosted in-house, and listed at /catalogue.

One move per shot, because rerolls are whole clips

A shot that pushes in, then tilts up, then holds, has three chances to go wrong and no way to fix one of them in isolation. Video bills per second of output, and a reroll buys the whole clip again, not the part you disliked. Two moves in one shot roughly doubles the surface area of a reroll while making the result harder to cut, because there is no clean frame to enter or leave on.

Split it. Push in as one shot, tilt up as another, hold as a third. Each one is shorter, each is independently rerollable, and the cut between them is yours to place. Ending each shot with a second of held framing gives an editor somewhere to make that cut without splicing mid-motion.

Shot 1: push in6 s
Shot 2: orbit left5 s
Shot 3: tilt up4 s
Shot 4: locked-off7 s

seconds of output per shot

An illustrative shot plan, not measured data: a 22-second sequence built from four single-move shots instead of one long take. Each bar is a clip you can reroll on its own, and RenderHour quotes a fixed price per run before you submit, so the plan is also the budget.

The short version

Describe the shot in one sentence. Name one move in grip language. Give that move a start framing and an end framing. Set the pace. Rule out shake, drift, zoom and cuts. Drive the look from a still whenever you have one, and keep each clip to a single idea so a reroll costs one shot instead of a sequence.

Copyable versions of these prompts, with the parts you are meant to change left in brackets, live in the prompt bank at /academy.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free →

Newsletter

Stay in the loop

Get the latest model comparisons and guides — no spam, unsubscribe anytime.