Camera control with Wan: the vocabulary that actually moves a camera
Flat generated footage is almost never a model problem. It is a direction problem. The prompt described a picture and left the camera to invent its own behaviour, so the camera invented drift. Here is the vocabulary that replaces drift with intent, the six moves worth learning, and the phrasing that makes each one land.
“Cinematic” is a review, not an instruction
The most common first attempt looks like this: cinematic shot of a lighthouse at dusk, dramatic, film look, 4k. What comes back is a lighthouse at dusk with a camera wandering gently around it, because nothing in that sentence says where the camera starts, where it ends, or how fast it travels between the two.
Every word in that prompt is about how you want to feel about the result. None of it is about what the camera does. Words like cinematic, epic, dramatic and beautiful are the note you give after a screening. They are not the note you give the operator before the take. Cut them entirely and spend the same space on geometry and pace, and the output changes more than any parameter you can tune.
A shot and a move are two different sentences
There are two things to describe, and they answer different questions. The shot is what is in frame: subject, framing, light, background, wardrobe. The move is what the camera does across the duration: direction, distance, speed.
Most prompts describe the shot in loving detail and say nothing about the move, then treat whatever motion appears as the model's fault. Motion always appears. Video output has a time axis and something has to fill it. When nothing constrains that axis you get the default: a slow lateral float, a breathing zoom, a handheld wobble nobody asked for. Drift is not a glitch. Drift is what motion looks like when it is unconstrained.
So write them as separate clauses, in order, and stop mixing them. Shot first, move second, constraints third.
- 01State the shotSubject, framing, light, location. One sentence. No camera language in it yet.
- 02Name one moveGrip language, not adjectives: push in, orbit left, tilt up. One move per shot.
- 03Fix the endpointsWhere the framing starts and where it ends. Wide to medium close-up, floor to skyline.
- 04Set the paceSteady walking pace, slow, one continuous take. Pace is what separates a push from a lunge.
- 05Rule out the restNo handheld shake, no zoom, no cut, no second subject. Short, camera-focused negatives.
The six moves worth knowing
Almost everything you will cut into a sequence is one of six moves, or a locked-off frame that deliberately refuses to be one. Learn the phrasing for these and you can direct the large majority of shots without reaching for anything exotic.
| Move | What to write | What goes wrong without it |
|---|---|---|
| Push in | A slow dolly push toward [subject], starting wide and ending at a medium close-up, steady walking pace. | A zoom instead of a dolly, so the background flattens rather than compressing, and the shot overshoots into a face-filling close-up. |
| Pull out | The camera pulls back from [subject] to reveal [context], ending on a wide that holds the whole [location] in frame. | The pull keeps going past the reveal and ends on an empty wide with the subject lost in it. |
| Orbit | Orbit the camera 90 degrees to the left around [subject], keeping it centred and the same size throughout. | A partial swing that stalls halfway, or an orbit that drifts closer as it travels and turns into a spiral. |
| Pan | The camera pans right from [starting element] to [ending element] on a fixed tripod head, no lateral travel. | The camera translates sideways instead of rotating, which reads as a track and breaks the parallax you wanted. |
| Tilt | The camera tilts up from [ground detail] to [upper element], pivoting in place, ending with [element] centred. | A crane move: the whole camera rises, so the low-angle geometry that made the tilt worth doing disappears. |
| Locked-off | Static locked-off camera on a tripod. Nothing moves except [one small motion]. No camera drift. | The default float. A shot that was supposed to be the calm anchor in the cut breathes through the whole take. |
Endpoints are what stop the overshoot
“Push in on the subject” names a direction and no destination. A direction with no destination runs until the clip ends, which is how you get shots that start nicely composed and finish on a nostril.
Naming the end framing turns an open-ended direction into a bounded one. Starting wide and ending at a medium close-up gives the move a finish line. So does ending with the horizon on the upper third, or ending back where it started for a full orbit. The vocabulary here is the standard framing ladder: wide, medium wide, medium, medium close-up, close-up, extreme close-up. Use two of those rungs per shot, one for the start and one for the end, and the distance travelled is decided before the render begins rather than during it.
The same trick fixes pace. A move with a named start and end has an implied speed, and adding steady or slow, no acceleration pins it down. Most shots that feel cheap are moving too fast for the distance they cover.
Negative constraints do more work than the description
Naming one move does not forbid the others. A push-in prompt can still come back with a push-in plus a handheld tremor plus a slow roll, because none of the extras were excluded. Four short negatives handle nearly all of it:
- No handheld shake. Removes the operator-breathing texture that gets added to anything vaguely documentary.
- No camera drift. The one that makes locked-off shots actually lock off.
- No zoom. Keeps a dolly a dolly, so the background compresses the way a physical push does.
- No cut, no scene change. Stops a single shot turning into a two-shot mini-edit halfway through the duration you paid for.
Keep negatives about camera behaviour, short, and few. A long list of things you do not want is another way of describing a picture, and it competes with the move for attention. Four is usually enough. If you find yourself writing eight, the shot is probably trying to do two things and wants splitting.
Let the image hold the art direction
Text-to-video asks one prompt to carry the look and the motion at once, and the two compete. Every extra clause about wardrobe, palette and lens is a clause that is not about the camera.
Image-to-video splits the job. A still you already approved carries the composition, the palette, the lighting direction and the character, so the prompt is free to be almost entirely about motion. Orbit the camera 90 degrees to the left around the subject, same size throughout, parallax reveals the doorway behind. No zoom, no cut. That is a complete instruction when the first frame has already settled everything else, and it is a much shorter thing to iterate on when the move is not quite right.
On RenderHour, Wan 2.7 ships text-to-video, image-to-video and audio-to-video variants, so the same family covers a still-driven shot and a text-driven one without changing your vocabulary between them. LTX-2.3 adds a video-to-video path when the input is already footage, and HunyuanVideo 2.0 covers text and image entry points as well. The whole open-weight catalogue is hosted in-house, and listed at /catalogue.
One move per shot, because rerolls are whole clips
A shot that pushes in, then tilts up, then holds, has three chances to go wrong and no way to fix one of them in isolation. Video bills per second of output, and a reroll buys the whole clip again, not the part you disliked. Two moves in one shot roughly doubles the surface area of a reroll while making the result harder to cut, because there is no clean frame to enter or leave on.
Split it. Push in as one shot, tilt up as another, hold as a third. Each one is shorter, each is independently rerollable, and the cut between them is yours to place. Ending each shot with a second of held framing gives an editor somewhere to make that cut without splicing mid-motion.
seconds of output per shot
The short version
Describe the shot in one sentence. Name one move in grip language. Give that move a start framing and an end framing. Set the pace. Rule out shake, drift, zoom and cuts. Drive the look from a still whenever you have one, and keep each clip to a single idea so a reroll costs one shot instead of a sequence.
Copyable versions of these prompts, with the parts you are meant to change left in brackets, live in the prompt bank at /academy.
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.