Video Scripts, Storyboards and AI-Assisted Production
11 min read · updated August 4, 2026
The time in a video does not sit where people assume. Machine assistance saves real hours in scripting, transcription, captioning and the rough cut, and it adds a review cost in generated footage that is usually larger than the generation time it replaced. The arithmetic is below.
Where the hours actually go
A short explanatory video — six to ten minutes, one presenter, screen recordings and some cutaways — decomposes into script, record, log, rough cut, fine cut, graphics, captions, and delivery. Script and fine cut dominate. Recording is short. Logging and captioning are pure mechanical time and are the least loved parts of the job, which is precisely why they are the ones worth automating first.
The question to ask of any tool here is not whether it produces something faster. It is whether the thing it produces can be accepted without being fully reviewed. Anything that must be watched end to end before use has a floor on its cost equal to its own duration, and that floor does not move however fast the generation gets.
Stage by stage: saved and spent
| Stage | Description |
|---|---|
| Script structure | Real saving. Turning a set of claims into a spoken-order outline is a structuring job. Write the claims yourself; the ordering pass is the same one used for prose. |
| Script prose | Mixed. Generated narration reads as written rather than spoken and has to be rewritten aloud anyway. Useful for a rough length estimate: at about 150 words per minute of delivery, word count is a runtime estimate. |
| Storyboard and shot list | Real saving on the list, not the boards. Extracting a shot list from a script is mechanical. Generated frames are useful as a mood reference and misleading as a composition reference, because what you can generate is not what you can shoot. |
| Transcription and logging | The largest single saving in the pipeline. See below. |
| Rough cut | Real saving where the edit is transcript-driven: cutting the text cuts the video. Does nothing for cuts driven by picture or performance. |
| Generated B-roll | Usually a net loss on anything that must match. Priced below. |
| Captions | Real saving with a mandatory correction pass. The machine produces the timing; a human fixes names, terms and line breaks. |
| Translated subtitles | Draft only. Register and idiom are exactly the things that break. |
| Thumbnails and titles | Generation is cheap and selection is the hard part, which is the same shape as the headline problem. |
Two rows point outward rather than being covered here. Subtitles into another language inherit every failure on the creative localisation page, and thumbnails and titles are the generate-wide, select-narrow problem with a picture attached, including the part about not being able to test your way to an answer at low volume.
The transcript is the leverage point
Almost every genuine saving in this list runs through one artefact. Once you have an accurate, time-coded transcript, a set of previously manual jobs become text operations.
- Logging. Finding the good take is a search over text rather than a scrub through footage.
- The paper edit. Deleting sentences in the transcript produces the rough cut, and it is far faster to argue about an edit in text than on a timeline.
- Filler removal. Mechanical, list-driven, safe.
- Chapters and timestamps. Structure over text with times attached.
- Captions and subtitles. The same object with line breaks applied.
- The written version. The article that accompanies the video is a rewrite of the transcript, not a new piece of work.
So the first investment is transcript accuracy, not editing tools. The quality bar and the arithmetic behind it are on the podcast production page, and everything there applies unchanged to video.
Generated footage and the review tax
Here is the calculation that decides whether generated B-roll is worth it. Assumptions are labelled and you should substitute your own.
GOAL: 10 minutes of usable cutaway footage. Assumptions (substitute your own): generation time 2 min of wall-clock per 1 min of output acceptance rate 40% of generated clips are usable in context review time 1.0× real time — you must watch it to judge it library alternative 35 min of searching, licensing and downloading To end up with 10 usable minutes at a 40% acceptance rate you must generate 25 minutes. generating 25 min output × 2 min/min = 50 min reviewing 25 min output × 1.0× real = 25 min ---------------------------------------------------- total 75 min stock library equivalent 35 min Break-even acceptance rate, holding everything else fixed: 10/a minutes generated × 3 min/min of handling = 35 min 30/a = 35 → a ≈ 0.86 So generated footage wins only if roughly 86% of what comes back is usable. Below that, the library is cheaper.
The insight is not the specific number, which depends entirely on your assumptions. It is the shape: review time does not fall as generation gets faster, so improvements in generation speed barely move the total. If a tool became instant, the 75 minutes above would become 25, still not far from the library’s 35.
Three situations flip the conclusion, and they are worth naming because they are where generated footage genuinely earns its place. When the shot does not exist in any library — your specific product, an abstract concept, a fictional scene — the alternative is not 35 minutes but a shoot. When the clip is very short, two seconds behind a title, review time collapses. And when a series needs a consistent look that a library cannot supply, consistency is worth paying review time for.
The cost nobody prices is continuity. Generated clips of the same subject differ in details a viewer notices without being able to say why — a different jacket, a room that has changed shape, a reflection that moves. Two clips of the same nominal scene cut together look wrong, and the fix is generating more and rejecting more, which drives the acceptance rate down exactly where you needed it up.
Captions: the quality bar in numbers
Auto-captions are a draft. The conventions below are widely used across broadcast and streaming subtitle specifications, and while the exact figures differ between houses, the shape is consistent enough to work from.
- Reading speed. Around 160–180 words per minute for adult audiences, or roughly 17 characters per second, is the common ceiling. Above it, viewers stop finishing lines.
- Line length. About 37–42 characters, two lines maximum on screen at once.
- Duration. Roughly one second minimum so a caption does not flash, and about seven seconds maximum so it does not outstay the shot.
- Line breaks fall at syntactic boundaries. Break after a clause, not between an article and its noun. This is the single most visible difference between machine and human captioning, and it is not something a reader can articulate — it just reads worse.
- Never break a caption across a shot change.
The correction pass therefore has a fixed shape: fix names and technical terms first, from a list you prepared; then re-break lines at clause boundaries; then check that no caption crosses a cut. Timing is usually the part you can accept as generated.
The rule that falls out of this
Everything above reduces to one test: automate the stages whose output you can verify by reading, and be sceptical of the stages whose output you must watch.
Transcripts, shot lists, chapter marks, captions and paper edits are all verifiable by reading, and reading is many times faster than real time. Generated footage, generated voice and generated music must be experienced at their own duration to be judged, so their review cost is bounded below by the length of the thing itself. That single asymmetry predicts which parts of a video pipeline machine assistance has already transformed and which parts it has quietly made more expensive.