folio 4 of 27
the ai video cloning blueprint: how to recreate a video with ai
- date
- 02 JUL 2025
- era
- pre-Fable
- status
- complete
- form
- essay
- updated
- 29 JUL 2026
how i turn a reference video into a shot brief, generation prompts, and an edit plan without confusing imitation for a finished film.
first published 2 jul 2025. refreshed 29 jul 2026. the first version was built around veo 3. the workflow survived the model names.
the video above is my 2025 reconstruction of what if tsar bomba hit new york?. i call this "cloning" as shorthand, but pixel-copying a creator's work is not the interesting part. the useful part is learning to see why a video works, then turning those choices into something you can test.
i wrote the first version of this guide with two good ideas buried under a giant prompt. here are the ideas that survived.
start with a shot brief, not a generation prompt
give the reference video to a multimodal model, then ask for a scene-by-scene brief. it is faster than pausing a video every two seconds and it gives you a document you can edit.
do not treat its first answer as ground truth. watch the reference with the brief open. correct the timestamps, call out the shots it missed, and separate what is visible from what the model guessed about mood or intent. this is where you decide what actually matters.
ask for these things:
- the video's premise and the rhythm of the edit
- a timestamped list of shots
- the subject, action, setting, light, palette, framing, and camera movement in each shot
- the soundscape: voice, music, ambience, and the moments where they change
- on-screen text and where it lands
- the job of each shot: establish, explain, surprise, release tension, or move the story forward
- anything uncertain, marked as uncertain rather than invented
this is the prompt i would use now:
watch this reference video and make a shot brief for a new video.
do not praise it or invent intent. separate observations from guesses.
start with the premise, overall visual language, edit rhythm, and audio approach.
then list every shot with:
- timestamp
- subject and action
- setting, light, and palette
- framing and camera movement
- sound, dialogue, music, and on-screen text
- the shot's job in the sequence
flag anything you cannot determine from the video.
turn every shot into a small, complete instruction
a generator does not know your edit. each prompt needs enough context to stand on its own. repeat the details that make the world recognizable: the character, wardrobe, place, time of day, lens language, palette, and audio texture. then give the clip one clear thing to do.
trying to make a whole scene happen in one prompt usually gives you a vague montage. break it into shots. the cut is where you get the rhythm back.
shot 04, intended for the moment after the flash
subject: a lone person in a dark jacket, seen from behind
action: stops at a storefront and looks up
setting: empty city street at blue hour, wet pavement, distant smoke
camera: slow push-in from a wide, eye-level shot
look: muted blue-grey palette, sodium streetlights, slight film grain
sound: wind, a far-off siren, no dialogue, no text on screen
continuity: keep the same jacket, street, palette, and grain as shots 01-03
the description is not magic. it is a checklist. if the result looks wrong, you can point to the decision that is wrong instead of vaguely asking the model to "make it better."
test one hard shot before making the whole thing
pick the shot that carries the visual identity of the piece. generate a few versions of that one first. it tells you whether the character description, visual language, and camera vocabulary are doing their job before you have fifty clips to throw away.
keep a tiny contact sheet or timeline of the candidates. choose with the reference beside it. this is not busywork. the model can make options, but it cannot know which imperfection is the one you want to keep.
use the model for the part it is actually good at
the tools changed since i first made this. google's current documentation points to gemini omni flash for general video generation and conversational editing. it keeps veo 3.1 for things such as native audio, video extension, and frame-specific generation. use the capability you need, not the model name from an old tutorial.
the current veo guide and google's prompt guide are worth checking before a big run. they cover the controls that tend to matter here: subject, action, scene, camera, visual style, audio, negative prompts, and frame-based workflows.
edit after generation
generation gives you material. the video happens in the edit.
put the clips on a normal timeline. compare the cuts to the reference. tighten the timing, rebuild the sound, add titles outside the model when the words need to be exact, and replace only the shots that break the sequence. the result can carry the same rhythm without pretending to be the original.
credit and consent
use a reference you made, licensed material, or something you have permission to study. do not use this to mimic a person's face or voice without consent, and do not present a generated study as if the original creator made it. the point is to learn visual language, then make your own work with it.
folio 4 of 27 · written in cusco, peru, 02 JUL 2025 · updated 29 JUL 2026