Two hundred clips get drafted cheaply first, reviewed by a person, and only the survivors are remade at full quality. That two-pass split is the craft in this job. Starting each clip from a real product photograph rather than a text description is what stops the product drifting between clips.
A small agency. One client, forty products, five vertical clips each for paid social. Eight seconds, product turning slowly, a different backdrop per clip. This is the third of four jobs in what AI can do that a chat window cannot, and it is the one where the interesting decision is about sequencing rather than about models.
The client will approve or reject these on one criterion: does the product look like the product. Everything else is negotiable.
Why the chat window stalls
Three separate walls, any one of which would be enough.
Most chat tools do not make video. The capability is simply absent. This is the fourth question in the guide's four-question test — does the tool you need exist in a chat box — and here the answer is no.
The ones that do are metered and capped. Per-clip pricing and a queue limit mean two hundred clips is two hundred sittings, spread across however many days the cap allows. The work is not hard; the process is just administratively impossible.
And they start from text. This is the one that matters. Describe the product in a prompt and the model builds it from that description each time, so it drifts — slightly wrong logo here, slightly wrong proportions there, differently in every clip. Two hundred clips of nearly-your-product is not a campaign. It is two hundred things the client rejects.
The models that do the work instead
Three models, and the order they run in is the actual craft.
LTX-Video does the draft pass. It is fast and cheap, so all two hundred concepts get made and looked at before any expensive time is spent on any of them.
Wan 2.2, image-to-video, does the finish pass on the survivors. It starts from the real product photograph and adds motion to it, so the product stays the product rather than being reinvented from a sentence. This is the direct answer to the drift problem.
RIFE fills in frames between the ones generated, turning a stuttery clip into smooth thirty-frames-a-second motion. Unglamorous, and the difference between footage that looks generated and footage that looks shot.
The run, step by step
- STEP 1Draft all 20050 min
- STEP 2You pick the keepers25 min
- STEP 3Final pass on 1403 hr 40 min
- STEP 4Smooth the motion35 min
- STEP 5Export verticals15 min
The final pass is nearly two-thirds of the run on its own, which is exactly why it happens second. Drafting all two hundred takes fifty minutes. Remaking one hundred and forty of them at full quality takes three and three-quarter hours. If you generated all two hundred at full quality you would spend roughly an extra hour and a half of machine time producing sixty clips that were always going to be thrown away.
There is one checkpoint in the middle, and it stays visible on purpose. The run pauses, shows you the drafts, and waits while you keep the good ones. Two minutes of your time buys most of the saving in this job.
What you will have when it is done
Finished vertical clips, one file per cut, ready to post. The rejected drafts, kept in case you change your mind about one. And a contact sheet of first frames, so you can find a specific clip without scrubbing through two hundred files looking for it.
The rejected drafts are not sentimentality. Concept review at draft quality is genuinely hard, and roughly one in ten rejections gets reversed once the campaign is assembled and something obvious is missing.
What the two minutes are actually for
The check-in is short, so it is worth being deliberate about what you are looking at. You are not judging quality — the drafts are meant to look rough, and a clip that looks poor at draft quality will usually look fine after the finish pass.
You are judging whether the idea works. Does the backdrop suit the product. Does the motion show the thing a buyer wants to see. Is the product recognisably itself, or has something gone wrong that no amount of finishing will fix.
Those are all questions a person answers in a second per clip and a model cannot answer at all, which is why this step exists and why it is the only human step in the middle of the run. Two hundred clips at a second each is three and a half minutes, and most people spend two.
A useful habit: reject on concept generously at this stage. The finish pass is where the cost is, and a clip you were lukewarm about at draft quality is rarely a clip you love at full quality.
The two-pass split, generalised
This is the third question in the model-selection framework: how many, and how good does each one need to be? The answer is rarely the same for every item.
Most large jobs split the same way this one does. A fast model runs over everything, a person looks, and a slower and better model runs over what survived. It is cheaper, it finishes sooner, and the work is better — not despite the human step but because of it. The person spends their two minutes on the decision only they can make, and no time at all on the execution.
The product photo job uses a quieter version of the same idea, sorting the difficult images out first so the expensive treatment goes only where it is needed.
Where this approach is weakest
Worth saying plainly. Generated video is still generated video, and eight seconds of a product turning on a backdrop is close to the best case for it — short, simple motion, a single subject, no people, no hands, no text on screen.
Push past that and the honest answer changes. Anything with a person in it, anything longer than about fifteen seconds, anything where the camera moves in a specific way: those still favour a shoot. This job works because the brief is narrow, and narrowing the brief was part of making it work.
How it works shows the same job as a handover, including where the check-in falls in the sequence.
