How to make a game trailer with AI, shot by shot
A trailer is not one long generation. It is six to ten short clips in a deliberate order, each doing one job, cut together over a music bed. Once you plan it that way, AI video stops being a slot machine and starts behaving like a render queue with a price on it.
- 01Key artOne frame that fixes the character, the palette, and the world.
- 02Shot listRows on paper: length, driving frame, and the job in the cut.
- 03Reference framesA still per location, every one carrying the same character.
- 04RenderImage to video, one beat per shot, a few takes each.
- 05CutPick takes, trim drifting tails, cut on motion.
- 06AudioMusic bed, one voice line if the trailer needs it.
Start from art you already own
Text to video invents the world again on every run. Ask it twice for your hero and you get two different characters wearing roughly similar armour. That is fine for a mood piece and fatal for a trailer, where the same face has to appear in five shots.
Image to video takes the composition, the palette, and the character design from a frame you control, then adds motion. If you already have key art, a splash screen, a marketing render, or a cleaned up in-engine screenshot, that is your input and most of your consistency problem is already solved. On the catalogue the relevant entries are Wan 2.7 for image to video when the shot has to look its best, HunyuanVideo 2.0 in the middle, and LTX-2.3 when you are burning through takes and want them cheap.
No key art yet? Make the still first. Images are billed per image and video per second of output, so ten attempts at a frame is a different order of spending from ten attempts at a clip. Iterate in stills, commit in video.
Write the shot list before you render anything
A trailer that works runs about twenty to thirty seconds. Write those seconds down as rows before you submit a single job. Each row gets a length, a driving frame, and one sentence saying what the shot is for. Rows you delete on paper cost nothing. Rows you discover halfway through a render session cost real money and, worse, tempt you into keeping a weak shot because you already paid for it.
| Shot | Length | Driven by | Job in the cut |
|---|---|---|---|
| 1. World establisher | 5s | Key art, wide | Sell the place. Slow camera, no character action. |
| 2. Hero reveal | 4s | Character reference frame | Sell the protagonist. One move, held long enough to read. |
| 3. Traversal | 3s | Second location still | Show movement and scale. First cut on action. |
| 4. Threat | 3s | Enemy still | Raise the stakes. Nothing else happens in this shot. |
| 5. Ability beat | 3s | Character frame, closer crop | The thing your game does that others do not. |
| 6. Logo card | 2s | Static plate | Title, date, platform. Held, not animated. |
That is twenty seconds of finished trailer across six shots. Note how much of the running time goes to the two shots at the front. Viewers decide in the first four seconds, so the establisher and the reveal earn their length while everything after them is a rhythm section.
The vocabulary that actually moves the camera
Video models respond to camera language, not to adjectives. Words like epic and cinematic do almost nothing. Words that describe an operator doing a thing do a lot: slow push in, dolly left, orbit around the subject, crane down, handheld follow, whip pan, rack focus, locked off static shot.
Two rules make those terms behave. First, one camera move per shot. A prompt asking for an orbit that becomes a push in gives you a model averaging two intentions, which reads as drift. Second, describe camera motion and subject motion as separate clauses, because the model will otherwise apply one to the other and swing the whole frame when you only wanted the character to turn.
Keep a fixed block of lens and lighting words in every prompt in the trailer: focal length, time of day, key light direction, grade. Changing that block between shots is a reliable way to make two clips of the same character look like two different games.
One reference frame, every shot
Character consistency comes from the input frame, not from the prompt. The working rule is simple: every shot featuring the hero is driven by the same reference frame, and what changes between shots is the camera and the motion, never the description of the character.
When you need the hero somewhere else, do not describe the new place to the video model. Build the new still first, carrying the character crop across into that background, and then animate the still. You are moving the consistency problem into image space, where a bad attempt costs one image instead of several seconds of video, and where you can look at the result before committing.
Two habits pay for themselves here. Keep a folder of approved reference frames with a filename per shot number, and keep the prompt for each shot in a text file next to it. When a shot needs a fourth take three days later, you want the exact input, not an approximation of it.
One beat per shot, for two separate reasons
The editing reason: a clip that lands, turns, and draws a sword is three chances to fail in one file. If any of the three goes wrong you rerun all of it. A clip that does one thing either works or fails cheaply, and it hands the editor a clean cut point at both ends.
The money reason: video is billed per second of output. Long clips are not just more expensive per attempt, they raise the cost of every rejection, because you throw away eight seconds to fix a problem that lived in one of them. Short shots keep your unit of waste small. Most catalogue video models cap a single run at a few seconds anyway, which is a constraint pushing you towards the right structure rather than away from it.
Cutting it, and where a still tail saves a rerun
Lay the music bed first and let it decide the cut points. Then place shots in the order on your list, cut on motion rather than on stillness, and resist trimming the establisher. Trailers rarely fail from being too slow at the start, they fail from having nothing to look at.
Generated clips tend to degrade at the end. Hands lose count, faces wobble, a background element slides. The fix is not another render. Trim the drifting tail, hold the last clean frame as a still for a few frames, and cut or dissolve out of that hold. On a beat, a two-frame freeze reads as an editorial choice rather than a fault, and it converts a failed take into a usable one.
If a character has to speak to camera, that is an audio to video job rather than a normal shot: Wan 2.7 and LTX-2.3 both take an audio track and drive the performance from it. Generate the line as speech, which is billed per character of input text, then drive the shot with it. And when the cut is assembled, Video summarization and Video Q&A are a decent sanity check: ask what happens in the trailer and see whether the answer matches the game you think you are selling.
The budget, per shot
Here is the part most guides skip. Because video is priced per second of output and the price is quoted before the job is submitted, you can cost the whole trailer on paper. Plan for three takes per shot, which is roughly what a well specified shot needs, and one take for the logo card.
rendered seconds and cost at a hypothetical $0.08/s
Fifty-six rendered seconds for twenty finished ones, at about $4.50 on that made up rate. Substitute the real per-second figure from the catalogue and the shape of the answer holds: the cost of a trailer is set by how many takes you plan, not by how long the trailer is. Two shots eat half the budget because they are the two long ones, which is also an argument for shooting them last, once the cheaper shots have taught you what the prompts need to say.
The failure mode that wrecks budgets is not an expensive model. It is rendering without a shot list, liking take seven of a shot you did not need, and discovering at the edit that the trailer has no threat beat.
Running it on RenderHour
RenderHour hosts its open-weight catalogue itself, so the models named above are all in the same place with the same billing model. Every job is a fixed price per run, quoted before you submit: video per second of output, images per image at a base resolution, speech per character of input text. If a request cannot be priced, it is refused rather than guessed at, which is the property that lets you plan a trailer as a spreadsheet instead of as a hope.
The minimum top-up is $10, which is enough to sit down and render a short trailer end to end. If you are producing them regularly, RH Pass starts at $20 a month. Browse the models and their prices on the catalogue and the plans on pricing.
Then go and write six rows on paper. The rendering is the easy part.
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.