Research · In this section

Visual Caption Treatment and Narration Alignment — Run 001

Status · REVISED

Research class

METHOD UPDATE

Completion date

July 27, 2026

Research question

Can Show Director turn a complete narrated short story into a readable long-form Story Video whose on-screen text follows the performance without obscuring the footage?

Prior hypothesis

Short, bold, two-line captions could improve the visual clarity of a complete narrated Story Video while preserving the accepted wording and existing production audio.

Method

One complete narrated short story was assembled with thirty approved owned video clips and its existing narration-and-music mix. The accepted text was divided into short display phrases without changing source words. An initial high-margin caption treatment was rendered and visually rejected. A corrected lower-third treatment was tested on a short real clip, then applied to the complete 32-minute production. Opening, midpoint, and ending frames were inspected; the complete video and audio streams were decoded; and browser load, play, pause, midpoint seek, and playback after refresh were verified.

Observations

  • The initial treatment produced a narrow vertical column of text that obscured the footage and was rejected.
  • Reducing the caption margins and using a restrained two-line lower-third restored readability and preserved the visual field.
  • The completed replacement runs for approximately 32 minutes 18 seconds at 1920 by 1080 with H.264 video and AAC audio.
  • The existing embedded narration-and-music mix remained unchanged; no second music bed or new speech render was introduced.
  • Browser playback loaded with a non-zero duration, sought to the midpoint, advanced during playback, paused, and remained playable after refresh.
  • The display transformation added line breaks only; the accepted source words remained unchanged.

Supporting evidence

  • The rejected render and its production failure reason remain recorded.
  • A real short-form style proof was inspected before the full replacement encode.
  • Opening, midpoint, and ending frames from the completed replacement show readable two-line captions without the earlier vertical-column failure.
  • Full video-stream and audio-stream decoding completed without errors.
  • The browser reported ready playback state, a 1,938.5-second duration, successful midpoint seeking, and advancing playback.
  • Thirty-three automated tests passed, including preservation of source words through the display-caption transformation.

Counterevidence

  • Caption timing is currently distributed proportionally from text length across the narration duration rather than derived from forced word or phrase alignment.
  • The proof does not establish that every caption remains synchronized with the spoken phrase throughout the complete work.
  • Only one long-form short story and one visual treatment were tested.
  • The owner has not yet watched the complete 32-minute video from beginning to end.

Conclusion

The revised two-line caption treatment materially improves readability and produces a technically valid, playable long-form Story Video without changing the accepted words or existing audio. The experiment does not establish precise voice-following synchronization because current cue timing remains proportional rather than forced-aligned.

Confidence

HIGH for visual repair, exact display wording, technical validity, and browser playback; LOW for phrase-level synchronization across the complete work

Limits

  • Precise word or phrase alignment has not yet been implemented or measured.
  • Complete human viewing and listening remains pending.
  • The result is based on one short story and one set of owned footage.
  • The proof establishes a usable visual method, not a general claim about every caption style or production.

Production consequence

Use the revised two-line treatment for private Story Video proofs. Do not treat proportional timing as exact voice-following synchronization or publish a captioned long-form work until alignment or full listening verifies that the text follows the performance.

Next unresolved question

Can local forced alignment produce phrase timings that remain synchronized across the complete narration without changing the accepted words or making captions visually frantic?

Support independent work

Help fund what comes next.

NOMOTO MEDIA publishes essays, investigations, fiction, audio, and films without a paywall. If the work is valuable to you, help support the next piece.