Home / Guides / What AI video production actually is

What AI video production actually is

There is a lot of noise about AI video. This is what the work looks like from inside it.

Generated, not filmed — MyThorneAI brand film Generated, not filmed — MyThorneAI brand film

The short version

AI video production means generating footage with a model instead of recording it with a camera. Everything around that — writing, storyboarding, editing, voice, sound, colour — stays roughly the same as it has always been.

That last point gets lost. People imagine you type a sentence and receive a finished commercial. What actually happens is that you replace one stage of a long pipeline, and the rest of the craft still has to be done by someone.

How a real project runs

1. Script

Nothing is generated first. A weak script produces a weak film regardless of how good the model is, and generative video is expensive enough in time that discovering your idea is thin at the render stage is a bad way to find out.

2. Storyboard

Shots get planned before they are generated. This is where you decide what the film actually looks like, and it is much cheaper to change here.

3. Generation

Individual shots are generated, usually several times each, and the best take is selected. This is closer to directing than to prompting — you are looking for a specific performance, not accepting the first plausible output.

4. Assembly

Edit, voiceover, sound design, colour grade. Ordinary post-production. This stage is what makes a set of clips feel like one film, and skipping it is the most common reason AI video looks like AI video.

You replace one stage of a long pipeline.
The rest of the craft still has to be done.

What the tools genuinely do well now

  • Physical scenes that would be expensive to stage — weather, crowds, large sets, impossible camera moves
  • Volume — producing forty episodes of a series without forty shoot days
  • Variants — once a concept exists, additional cuts and aspect ratios are cheap
  • Sensitive subjects — animated formats that avoid depicting real, identifiable people

Where it still falls over

Being straight about this is more useful than pretending otherwise.

  • Specific real people. If your campaign needs an actual named person, you need a camera.
  • Long unbroken takes. Coherence over long durations remains unreliable.
  • Precise text on screen. Generated text is still unreliable, so on-screen typography is usually added in post rather than generated.
  • Hands, and fine physical interaction. Better than a year ago, still worth shot-planning around.
  • Documentary truth. Anything that needs to be evidence of a real event is out of scope by definition.

Which tools

Google Veo 3, Kling and Seedance are the current mainstays for generation. They have different strengths per shot, so the sensible approach is to choose per shot rather than commit to one.

This list will be out of date within months. That is normal, and it is why the durable part of the skill is the writing and the edit, not the tool.

Is it right for your project?

Reasonable fit: brand films, explainers, commercials, educational series, product and social ads, concept and awareness films.

Poor fit: anything requiring a specific real individual, live event coverage, or documentary evidence.

If you are unsure, describing the project to someone who does this daily takes five minutes and will save you a lot of guessing.

Need this made?

I produce the thing this article describes. Tell me what you need.