# 200 short videos of a product that has to look right

The client approves or rejects on one criterion: does the product look like the product. Everything else is negotiable, including the schedule.

> **The short answer.** Two hundred clips get drafted cheaply first, reviewed by a person, and only the survivors are remade at full quality. That two-pass split is the craft in this job. Starting each clip from a real product photograph rather than a text description is what stops the product drifting between clips.

A small agency. One client, forty products, five vertical clips each for paid social. Eight seconds, product turning slowly, a different backdrop per clip. This is the third of four jobs in [what AI can do that a chat window cannot](https://gpuvault.io/guides/asking-vs-running/), and it is the one where the interesting decision is about sequencing rather than about models.

The client will approve or reject these on one criterion: does the product look like the product. Everything else is negotiable.

## Why the chat window stalls

Three separate walls, any one of which would be enough.

**Most chat tools do not make video.** The capability is simply absent. This is the fourth question in the guide's four-question test — does the tool you need exist in a chat box — and here the answer is no.

**The ones that do are metered and capped.** Per-clip pricing and a queue limit mean two hundred clips is two hundred sittings, spread across however many days the cap allows. The work is not hard; the process is just administratively impossible.

**And they start from text.** This is the one that matters. Describe the product in a prompt and the model builds it from that description each time, so it drifts — slightly wrong logo here, slightly wrong proportions there, differently in every clip. Two hundred clips of nearly-your-product is not a campaign. It is two hundred things the client rejects.

## The models that do the work instead

Three models, and the order they run in is the actual craft.

**LTX-Video** does the draft pass. It is fast and cheap, so all two hundred concepts get made and looked at before any expensive time is spent on any of them.

**Wan 2.2**, image-to-video, does the finish pass on the survivors. It starts from the real product photograph and adds motion to it, so the product stays the product rather than being reinvented from a sentence. This is the direct answer to the drift problem.

**RIFE** fills in frames between the ones generated, turning a stuttery clip into smooth thirty-frames-a-second motion. Unglamorous, and the difference between footage that looks generated and footage that looks shot.

## The run, step by step

**The run: 200 short videos of a product that has to look right.** $18.60 against a $25.00 cap.

| Step | How long |
|---|---|
| Draft all 200 | 50 min |
| You pick the keepers | 25 min |
| Final pass on 140 | 3 hr 40 min |
| Smooth the motion | 35 min |
| Export verticals | 15 min |

Machine time 5h 45m; about ~40 min of your attention.

The final pass is nearly two-thirds of the run on its own, which is exactly why it happens second. Drafting all two hundred takes fifty minutes. Remaking one hundred and forty of them at full quality takes three and three-quarter hours. If you generated all two hundred at full quality you would spend roughly an extra hour and a half of machine time producing sixty clips that were always going to be thrown away.

There is one checkpoint in the middle, and it stays visible on purpose. The run pauses, shows you the drafts, and waits while you keep the good ones. Two minutes of your time buys most of the saving in this job.

## What you will have when it is done

Finished vertical clips, one file per cut, ready to post. The rejected drafts, kept in case you change your mind about one. And a contact sheet of first frames, so you can find a specific clip without scrubbing through two hundred files looking for it.

The rejected drafts are not sentimentality. Concept review at draft quality is genuinely hard, and roughly one in ten rejections gets reversed once the campaign is assembled and something obvious is missing.

## What the two minutes are actually for

The check-in is short, so it is worth being deliberate about what you are looking at. You are not judging quality — the drafts are meant to look rough, and a clip that looks poor at draft quality will usually look fine after the finish pass.

You are judging whether the idea works. Does the backdrop suit the product. Does the motion show the thing a buyer wants to see. Is the product recognisably itself, or has something gone wrong that no amount of finishing will fix.

Those are all questions a person answers in a second per clip and a model cannot answer at all, which is why this step exists and why it is the only human step in the middle of the run. Two hundred clips at a second each is three and a half minutes, and most people spend two.

A useful habit: reject on concept generously at this stage. The finish pass is where the cost is, and a clip you were lukewarm about at draft quality is rarely a clip you love at full quality.

## The two-pass split, generalised

This is the third question in the [model-selection framework](https://gpuvault.io/guides/asking-vs-running/): how many, and how good does each one need to be? The answer is rarely the same for every item.

Most large jobs split the same way this one does. A fast model runs over everything, a person looks, and a slower and better model runs over what survived. It is cheaper, it finishes sooner, and the work is better — not despite the human step but because of it. The person spends their two minutes on the decision only they can make, and no time at all on the execution.

The [product photo job](https://gpuvault.io/guides/asking-vs-running/product-photos/) uses a quieter version of the same idea, sorting the difficult images out first so the expensive treatment goes only where it is needed.

Run one of these:

- [LTX-Video](https://gpuvault.io/model/ltx-video/) — Generates short video clips from a description or a still image. From $0.89/hr.
- [Wan 2.1](https://gpuvault.io/model/wan-2-1/) — The best-looking open video model. Slow and worth it for finished work. From $2.25/hr.
- [RIFE](https://gpuvault.io/model/rife/) — Makes video smoother, or turns normal footage into clean slow motion. From $0.45/hr.

## Where this approach is weakest

Worth saying plainly. Generated video is still generated video, and eight seconds of a product turning on a backdrop is close to the best case for it — short, simple motion, a single subject, no people, no hands, no text on screen.

Push past that and the honest answer changes. Anything with a person in it, anything longer than about fifteen seconds, anything where the camera moves in a specific way: those still favour a shoot. This job works because the brief is narrow, and narrowing the brief was part of making it work.

[How it works](https://gpuvault.io/how-it-works/) shows the same job as a handover, including where the check-in falls in the sequence.

Part of our guide to [What AI can do for your business that a chat window cannot](https://gpuvault.io/guides/asking-vs-running/).

## Questions people ask

### Why start from a photo instead of a description?

Because a text description is not precise enough to hold a product steady across two hundred clips. Describe it and each clip drifts a little differently — the logo moves, the proportions shift, the colour is not quite right. Starting from the real photograph anchors every clip to the actual product.

### What is the check-in partway through?

About two minutes of your time. The run pauses after drafting all two hundred clips and shows you the rough versions so you can keep the ones that work. Only the keepers get remade at full quality, which is where most of the cost and nearly all of the time goes.

### Why not just generate everything at full quality?

Because roughly a third of clips get rejected on concept rather than execution, and full-quality generation is the expensive step. Making two hundred drafts and one hundred and forty finals costs a fraction of two hundred finals, finishes sooner, and produces better work because a person chose.

### Can chat tools not do this?

Most do not make video at all. The ones that do meter you per clip and cap how many you can queue, so two hundred clips means two hundred separate sittings. They also start from text rather than from your photograph, which reintroduces the drift problem this whole approach exists to avoid.

### Is the eighteen-dollar figure real?

It is illustrative, modelled on two hundred eight-second vertical clips rather than measured on our own runs. Clip length, resolution and how many survive the draft review all move it substantially. Every session shows the price and a spend cap before anything starts.

---

*Published August 24, 2026. Source: https://gpuvault.io/guides/asking-vs-running/short-video/ — GPUVault, GPU rental by the minute.*
