Skip to content

Video Scripts, Storyboards and AI-Assisted Production

11 min read · updated August 4, 2026

The time in a video does not sit where people assume. Machine assistance saves real hours in scripting, transcription, captioning and the rough cut, and it adds a review cost in generated footage that is usually larger than the generation time it replaced. The arithmetic is below.

Where the hours actually go

A short explanatory video — six to ten minutes, one presenter, screen recordings and some cutaways — decomposes into script, record, log, rough cut, fine cut, graphics, captions, and delivery. Script and fine cut dominate. Recording is short. Logging and captioning are pure mechanical time and are the least loved parts of the job, which is precisely why they are the ones worth automating first.

The question to ask of any tool here is not whether it produces something faster. It is whether the thing it produces can be accepted without being fully reviewed. Anything that must be watched end to end before use has a floor on its cost equal to its own duration, and that floor does not move however fast the generation gets.

Stage by stage: saved and spent

StageDescription
Script structureReal saving. Turning a set of claims into a spoken-order outline is a structuring job. Write the claims yourself; the ordering pass is the same one used for prose.
Script proseMixed. Generated narration reads as written rather than spoken and has to be rewritten aloud anyway. Useful for a rough length estimate: at about 150 words per minute of delivery, word count is a runtime estimate.
Storyboard and shot listReal saving on the list, not the boards. Extracting a shot list from a script is mechanical. Generated frames are useful as a mood reference and misleading as a composition reference, because what you can generate is not what you can shoot.
Transcription and loggingThe largest single saving in the pipeline. See below.
Rough cutReal saving where the edit is transcript-driven: cutting the text cuts the video. Does nothing for cuts driven by picture or performance.
Generated B-rollUsually a net loss on anything that must match. Priced below.
CaptionsReal saving with a mandatory correction pass. The machine produces the timing; a human fixes names, terms and line breaks.
Translated subtitlesDraft only. Register and idiom are exactly the things that break.
Thumbnails and titlesGeneration is cheap and selection is the hard part, which is the same shape as the headline problem.

Two rows point outward rather than being covered here. Subtitles into another language inherit every failure on the creative localisation page, and thumbnails and titles are the generate-wide, select-narrow problem with a picture attached, including the part about not being able to test your way to an answer at low volume.

The transcript is the leverage point

Almost every genuine saving in this list runs through one artefact. Once you have an accurate, time-coded transcript, a set of previously manual jobs become text operations.

  • Logging. Finding the good take is a search over text rather than a scrub through footage.
  • The paper edit. Deleting sentences in the transcript produces the rough cut, and it is far faster to argue about an edit in text than on a timeline.
  • Filler removal. Mechanical, list-driven, safe.
  • Chapters and timestamps. Structure over text with times attached.
  • Captions and subtitles. The same object with line breaks applied.
  • The written version. The article that accompanies the video is a rewrite of the transcript, not a new piece of work.

So the first investment is transcript accuracy, not editing tools. The quality bar and the arithmetic behind it are on the podcast production page, and everything there applies unchanged to video.

Generated footage and the review tax

Here is the calculation that decides whether generated B-roll is worth it. Assumptions are labelled and you should substitute your own.

GOAL: 10 minutes of usable cutaway footage.

Assumptions (substitute your own):
  generation time      2 min of wall-clock per 1 min of output
  acceptance rate      40% of generated clips are usable in context
  review time          1.0× real time — you must watch it to judge it
  library alternative  35 min of searching, licensing and downloading

To end up with 10 usable minutes at a 40% acceptance rate you must
generate 25 minutes.

  generating   25 min output × 2 min/min   = 50 min
  reviewing    25 min output × 1.0× real   = 25 min
  ----------------------------------------------------
  total                                      75 min

  stock library equivalent                   35 min

Break-even acceptance rate, holding everything else fixed:
  10/a minutes generated × 3 min/min of handling = 35 min
  30/a = 35  →  a ≈ 0.86

So generated footage wins only if roughly 86% of what comes back is
usable. Below that, the library is cheaper.

The insight is not the specific number, which depends entirely on your assumptions. It is the shape: review time does not fall as generation gets faster, so improvements in generation speed barely move the total. If a tool became instant, the 75 minutes above would become 25, still not far from the library’s 35.

Three situations flip the conclusion, and they are worth naming because they are where generated footage genuinely earns its place. When the shot does not exist in any library — your specific product, an abstract concept, a fictional scene — the alternative is not 35 minutes but a shoot. When the clip is very short, two seconds behind a title, review time collapses. And when a series needs a consistent look that a library cannot supply, consistency is worth paying review time for.

The cost nobody prices is continuity. Generated clips of the same subject differ in details a viewer notices without being able to say why — a different jacket, a room that has changed shape, a reflection that moves. Two clips of the same nominal scene cut together look wrong, and the fix is generating more and rejecting more, which drives the acceptance rate down exactly where you needed it up.

Captions: the quality bar in numbers

Auto-captions are a draft. The conventions below are widely used across broadcast and streaming subtitle specifications, and while the exact figures differ between houses, the shape is consistent enough to work from.

  • Reading speed. Around 160–180 words per minute for adult audiences, or roughly 17 characters per second, is the common ceiling. Above it, viewers stop finishing lines.
  • Line length. About 37–42 characters, two lines maximum on screen at once.
  • Duration. Roughly one second minimum so a caption does not flash, and about seven seconds maximum so it does not outstay the shot.
  • Line breaks fall at syntactic boundaries. Break after a clause, not between an article and its noun. This is the single most visible difference between machine and human captioning, and it is not something a reader can articulate — it just reads worse.
  • Never break a caption across a shot change.

The correction pass therefore has a fixed shape: fix names and technical terms first, from a list you prepared; then re-break lines at clause boundaries; then check that no caption crosses a cut. Timing is usually the part you can accept as generated.

Where captions are a legal accessibility requirement, the applicable standard sets the bar rather than any convention here, and machine output alone is generally not accepted as compliant. Check the standard that applies to your jurisdiction and platform.

The rule that falls out of this

Everything above reduces to one test: automate the stages whose output you can verify by reading, and be sceptical of the stages whose output you must watch.

Transcripts, shot lists, chapter marks, captions and paper edits are all verifiable by reading, and reading is many times faster than real time. Generated footage, generated voice and generated music must be experienced at their own duration to be judged, so their review cost is bounded below by the length of the thing itself. That single asymmetry predicts which parts of a video pipeline machine assistance has already transformed and which parts it has quietly made more expensive.