What you will build
This guide does not start by filling a video with sentences and image ideas. You will first decide how far each scene should move the viewer's understanding. Then you will assign voice, captions, and visuals to that job.
By the end, you will have:
- one change between the viewer's starting and ending state;
- a distinct job for every scene;
- a question that each scene hands to the next;
- a storyboard table for VMAI input and media review.
This is a pre-production worksheet, not a click-by-click interface tutorial. You can complete it without an account or a live project screen.
Step 1: Write the starting and ending states
A storyboard is a map of change, not merely a shot list. Write the before and after in one line each.
textBefore watching: The viewer ____________________. After watching: The viewer ____________________.
Example:
textBefore watching: The viewer assumes a weak automatic subtitle draft must be generated again from scratch. After watching: The viewer knows to adjust language, context, and line length, then edit only what remains wrong.
If the gap is too large, one video probably cannot deliver it.
textBefore watching: The viewer knows nothing about video production. After watching: The viewer creates a perfect professional video.
That goal tries to fit several rounds of learning and production into one piece. Reduce it to a first decision.
textBefore watching: The viewer does not know which source to use. After watching: The viewer chooses Idea, Script, Audio, or Subtitle based on the material already available.
Step 2: List the content first, then label its job
Do not set the order yet. Write one line for every point you are tempted to include.
For a video about storyboards, the raw list might look like this:
- Every scene needs a job.
- The hook matters.
- Repeating a claim makes a video feel slow.
- Visuals should carry different information.
- The ending needs a next step.
- Extra scenes should be merged.
- Captions should be short.
Now label what each line contributes.
| Content | Possible job |
|---|---|
| Repeating a claim makes a video feel slow | Help the viewer recognize the problem |
| Every scene needs a job | Explain the decision rule |
| Visuals should carry different information | Apply the rule to visual choices |
| Extra scenes should be merged | Give the viewer an editing action |
| The ending needs a next step | Leave an action after the video |
"The hook matters" and "Captions should be short" may be true without being necessary for this video's promised change. Move lines with no clear job into a parking lot for another piece.
Step 3: Give every scene one viewer question
Once the jobs are visible, write the question each scene must answer.
textQuestion for Scene 1: Question for Scene 2: Question for Scene 3:
Useful questions create a handoff.
| Scene | Viewer question | Scene job |
|---|---|---|
| 1 | Why can a short video still feel slow? | Help the viewer recognize the problem |
| 2 | How do I find the repetition? | Explain the diagnostic rule |
| 3 | Should I delete every repeated scene? | Compare merging with changing the job |
| 4 | How should the visuals differ? | Apply distinct visual functions |
| 5 | What should I write before input? | Leave an immediate planning action |
If Scenes 2 and 3 answer the same question, merge them or sharpen the difference. If the questions jump too far apart, a short bridge may be necessary.
A bridge does not introduce another topic. It shows why the last answer creates the next question.
Step 4: Check neighboring scenes for repetition
Similar scenes are harder to spot when you stare at the whole board. Compare them in pairs.
The sentences repeat
Reduce each scene to one line.
textScene A: Rework increases when there is no plan. Scene B: Without preparation, you have to rebuild the video.
Both reach the same conclusion. They are merge candidates.
The evidence repeats
One scene can make the claim; the next should help the viewer test it. If both scenes only assert the idea, change the second into an example, comparison, number, or real procedure.
The visuals repeat
Different files can perform the same function.
textScene A: A person frustrated at a laptop. Scene B: A group frustrated in a meeting.
Both communicate only "there is a problem." Change the second visual to a sequence, comparison, result, or next action.
The energy repeats
Check sentence length, delivery speed, camera motion, and music intensity. Hold briefly on a decision. Slow down for a concrete example. Contrast gives the viewer a sense of progression.
Step 5: Give the visual a separate function
A content label such as "problem scene" does not tell you what the screen should carry. Add an information function for the visual.
| Visual function | What it should show | What to avoid |
|---|---|---|
| Subject | The person, product, or situation | Decorative imagery with no context |
| Evidence | A result, record, comparison, or visible change | Generic success imagery |
| Process | Sequence, movement, or transformation | A dense screen whose steps cannot be read |
| Contrast | A/B, before/after, wrong/improved | Two images with no visible difference |
| Destination | The next page, worksheet, or result | Several calls to action on one screen |
Pair the content job with the visual function.
textContent job: Explain how to diagnose repetition. Visual function: Contrast Visual plan: Place two scenes with the same claim beside a revised claim-to-evidence pair.
That plan gives you a better basis for image search or source media selection than a broad mood keyword.
Step 6: Keep voice, caption, and visual from repeating one sentence
The three layers should support the same purpose. They do not need to copy the same line.
| Layer | Job |
|---|---|
| Voice | Explain the reason and context |
| Caption | Preserve the phrase or conclusion worth remembering |
| Visual | Show the subject, process, or difference that words would describe poorly |
Weak layout:
textVoice: Merge scenes that have the same job. Caption: Merge scenes that have the same job. Visual: Large text saying "Merge scenes that have the same job."
More useful layout:
textVoice: If two neighboring scenes reach the same conclusion, they are merge candidates even when the wording changes. Caption: Same conclusion? Merge. Visual: Two duplicate storyboard cards combining into one.
If the video has no captions or will often play muted, redistribute the work for that viewing situation.
Step 7: Set time by reading and comprehension, not equal scene slots
Equal scene lengths can make the pacing feel mechanical. Estimate each scene in this order:
- Read the voice line naturally.
- Shorten the caption until it can be read comfortably.
- Leave enough time to notice the comparison or visual change.
- Keep only the pause required before the handoff.
Use a range and an exit condition rather than a precise number that looks more certain than it is.
textEstimated range: 5–7 seconds Exit condition: The viewer can see the comparison once and read the main caption.
If the scene stays after the idea is clear, shorten it. If it leaves before the idea lands, reduce the information or split the scene.
Step 8: Move the storyboard into the VMAI input mode
VMAI starts new videos from four input modes: Idea, Script, Audio, and Subtitle. The same storyboard works for all four, but the transfer step changes.
Starting with Idea
Describe the jobs and progression rather than demanding an exact scene count.
textAudience: creators whose short AI videos still feel slow Change after watching: diagnose repeated scene jobs before cutting seconds Flow: - recognize repetition in their own video - diagnose repeated claims, visuals, and energy - compare deletion with changing a scene's job - assign one job to each scene Ending: write the first three storyboard rows
Starting with Script
Leave a job note between sections, then write the voice line.
text[Scene job: Recognize the problem] If a 60-second video still feels slow, check the scene jobs before the runtime. [Scene job: Learn the diagnostic] When neighboring scenes reach the same conclusion, they repeat even if the wording changes.
Starting with Audio
Label each paragraph in the recording script before you record. If the audio already exists, mark passages with the same job in the transcript before deciding on edit points. Keep the source file intact and record the intended edits separately.
Starting with Subtitle
Do not assume every SRT or VTT cue is a story scene. Caption display units and narrative scene units are different. Several cues may support one scene job, while one long cue may need to be separated across two jobs.
Step 9: Review scenes in groups of three
A paused scene can reveal image quality, but it rarely reveals story repetition. Watch the previous, current, and next scenes together.
- The current scene answers a different question from the previous one.
- The visual carries a new type of information.
- Voice, caption, and visual do not copy one sentence.
- A question remains that makes the next scene useful.
- Sentence length and visual motion do not repeat the same pattern throughout.
- Removing a scene would change the viewer's understanding.
- Important evidence or conditions are still present.
- The visual plan contains no account details, emails, project names, customer data, or other private information.
When you find repetition, do not delete on instinct. First choose whether to merge it, turn it into evidence, turn it into a comparison, or turn it into the next action.
Copy-ready AI video storyboard
textAI Video Storyboard 1. Before and after - Viewer state before watching: - Viewer state after watching: - One change this video will deliver: 2. Parking lot - Content that is not required for this video: - Content to move into another video: 3. Scene row Scene number: - Current viewer question: - Answer in this scene: - Scene job: - Main voice line: - Short caption: - Visual function (subject/evidence/process/contrast/destination): - Visual plan: - Question handed to the next scene: - Estimated time range: - Exit condition: - Keep/merge/change job/delete: 4. Neighbor repetition check - Do two scenes reach the same conclusion? - Are claims repeating without evidence? - Do neighboring visuals have the same function? - Do energy and sentence length stay unchanged? - Does the previous scene make the next one necessary? 5. VMAI input prep - Input mode (Idea/Script/Audio/Subtitle): - Scene jobs to preserve in the input: - Changes required before upload: - Three-scene groups to inspect during media review:
The goal is not to fill every field with long notes. The board is finished when you can quickly explain why each scene exists and how it differs from the scenes around it.
For the strategy behind this worksheet, read Your AI Video Isn't Too Long. Its Scenes Are Repeating Themselves. If the overall assignment is still unclear, start with Create a One-Page AI Video Production Brief.
Tutorial Complete!
You've read this tutorial. Move on to the next step.
Found this tutorial helpful?
Subscribe to get notified when we publish new tutorials and guides. No spam, unsubscribe anytime.
Comments
Please login to leave a comment
LoginNo comments yet. Be the first to comment!