shortshort
GUIDE

How are word-by-word captions timed, and which caption style should I use?

How word-by-word captions are timed from a transcript, how words are grouped on screen, and how to choose between four caption styles for vertical shorts.

The short answer

Word-by-word captions give every spoken word its own start and end time, so captions can follow speech one word at a time. shortshort gets those times from OpenAI's whisper-1, the model OpenAI names for word timestamps as of September 2026, then groups words into short pages. Pick Spotlight for talking heads, Impact for one claim, Karaoke for explanations, Minimal when the picture matters most.

Sources checked on 22 September 2026, each listed at the end of this guide with its date.

Where the timing comes from: one time per word

Word-by-word captions need a transcript in which every word has its own start and end time. OpenAI's speech-to-text guide names whisper-1 as the model to use for word or segment timestamps, and the timestamp_granularities[] parameter only works with whisper-1 (checked September 2026). shortshort first analyses the whole video to choose the shorts, then asks whisper-1 for word-level timestamps on the passages it kept.

Once the shorts are chosen, each short's selected passages are cut out as 16 kHz mono audio and transcribed with whisper-1. The last passage runs on to the longest ending the short is allowed, plus 2 s, to see where the sentence finishes. The word times are then mapped back onto the source timeline, so each short gets timing measured on its own audio.

Some times cannot be measured. A Whisper word whose end is not after its start gets a window of up to 200 ms and is flagged as estimated. In the editor, those words show a ≈ sign next to their timestamp.

How words are grouped into pages on screen

Words are shown in small groups called pages. A new page starts when the current page already has 4 words, when the next word would push the text past 34 characters, when the next word starts more than 450 ms after the page ends, or at a clip boundary. Words that start at the same instant stay on the same page.

A page appears at its first word's start time and disappears at its last word's end time, trimmed so it never overlaps the next page. There is no minimum on-screen time, so captions disappear during pauses longer than 450 ms.

DCMP's Captioning Key, written for educational media, says captions should stay up long enough to be read completely and stay in sync with the audio. It asks for at least 40 frames (1 s 10 frames) and at most 6 s per caption, and no more than two lines. A page of one or two fast words can stay up for less than that, so watch the preview when the speaker talks in short bursts.

The four styles, side by side

Every short stores one style and an on/off switch for captions. Spotlight is the default. The style changes what happens to the word being spoken and to emphasized words.

Emphasized words are picked by the AI, which is asked for 3 to 5 per 30 s; they stay at least 2 s apart, with a cap per short. Words with estimated timing are never picked, and the text is never changed. You can also toggle emphasis on any word yourself. Because emphasized words are at least 2 s apart, many Impact pages show no highlight at all.

StyleTextWord being spokenEmphasized wordsshortshort suggests
Spotlight (default)Uppercase, white with a black outline. Size set from the longest word on the page (34 to 74 px), so the line does not reflowSits on a block in the highlight colour. No animation, no size changeIgnoredTalking heads
ImpactUppercase, 64 pxNot highlightedUp to 86 px (smaller for long words), highlight colour, then a highlight block, or an underline for underline emphasis, with a pop (up to 1.26x, -3 degree tilt, 12 px lift, fading over 360 ms)A clip built on one number or claim
KaraokeMixed case, 58 px. The whole page stays visibleTurns the highlight colourHighlight colour with the pop animationExplanations
MinimalMixed case, weight 500, each word on a dark box (#111 at about 87% opacity)Not highlightedNo effectWhen the picture matters more than the words

How to pick a style, a colour and a font

Start from what the viewer should look at. A speaker on camera fits Spotlight. A clip that hangs on one figure or one strong sentence fits Impact. A course or webinar passage that walks through steps fits Karaoke: the whole page stays readable while the spoken word lights up. Slides, demos and wide shots fit Minimal.

Case and background matter too. DCMP prefers mixed case for readability and keeps capitals for shouting: Spotlight and Impact are all caps, Karaoke and Minimal are not. DCMP also prefers white text with a shadow, and a translucent box behind it, especially on light backgrounds, which is close to Spotlight's outlined white words and Minimal's dark boxes.

The highlight colour can be any #rrggbb value, and the ink drawn on it switches between dark and light using the WCAG relative luminance formula. The default #dcff65 gives about 16:1 contrast with dark ink. WCAG 2.2 asks for 4.5:1 for text and 3:1 for large text on web content; its Understanding page does not mention video captions, but it is a useful reference. With some mid-tone colours the automatic ink falls below 3:1, so check contrast when you leave the default.

  • The caption font menu has 14 embedded fonts. The default is Grotesque (Inter).
  • Rendering waits for the chosen font to load, and fails rather than export in a fallback font.
  • Colour and font can be changed per video on any plan. The Studio plan saves them as one brand preset, preselected on every new video.
  • In Karaoke, the highlight colour is the text colour itself, drawn over the picture, so very dark colours read poorly.

Placement and proofreading before you render

On the 1080 x 1920 frame, the caption block sits 70 px from the left edge, 100 px from the right edge and 330 px above the bottom. It is centred, wraps onto several lines if needed, and has a drop shadow. WCAG guidance says captions should not obscure relevant information, and DCMP says they should not cover speakers' names, faces or mouths. Check the preview for shots where the block lands on a face or on slide text.

Proofread the words before rendering. W3C's accessibility guidance says automatically generated captions do not meet accessibility requirements unless they are confirmed to be fully accurate, and OpenAI notes that Whisper accuracy varies by language. The caption editor lists every word with its timestamp: retype a word (up to 80 characters), clear it to remove it, or toggle emphasis. Names, numbers and jargon are the usual suspects.

If the transcript is unusable, import an SRT file (up to 2 MB) timed to the source video. Each cue's time is split evenly across its words, so word timing is approximate and you should check the synchronization in the preview. The preview plays the same composition that renders the MP4, so what you see is what gets exported.

Burned-in captions and platform captions

shortshort burns captions into the picture, which makes them open captions: W3C describes open captions as always displayed and impossible to turn off. They also carry spoken words only. W3C and WCAG define captions as including non-speech audio such as sound effects, music, laughter and speaker identification, so a speech-only track should not be presented as full accessibility captions.

YouTube and TikTok can both add their own automatic captions, so the same words may appear twice on screen. Neither platform page checked says what happens with burned-in captions. After upload, look at the post and edit or remove the platform captions if they duplicate yours. For a separate closed-caption track, 'Export captions' under 'Edit files' saves lumo-captions.srt, built from the same pages and timed on the short's own timeline. YouTube accepts basic SRT files without style markup, so the four styles do not carry over.

PlatformAutomatic captionsWhat the creator can doCaption filesSource (checked 22 Sep 2026)
YouTube ShortsGenerated on long-form videos and Shorts, published automatically when availableReview them: YouTube warns they can misrepresent speech (accents, dialects, background noise)Basic .srt accepted, no style markupYouTube Help
TikTokGenerated automatically for uploaded videos; the creator picks the caption languageEdit or remove them after posting; creator captions can be styled (font style, colour). Viewers can switch auto captions on or offNot checkedTikTok Support

With shortshort

shortshort analyses the whole video to choose the shorts, then transcribes each short's passages with whisper-1 at word level and maps the times back onto the source. Captions are on by default in Spotlight, on every plan including Free. Style, highlight colour and font (14 fonts) can be changed per short. Captions stay in the spoken language, carry spoken words only, and vanish during pauses over 450 ms.

Questions people ask

Can viewers turn off captions that are burned into a short?

No. Burned-in captions are open captions, which W3C describes as always displayed and impossible to turn off. In shortshort, 'Download all' puts each captioned short in the ZIP twice, with and without captions, so you can post the clean version with a separate caption track.

Are word-by-word captions from shortshort accessibility captions?

Not on their own. W3C defines captions as covering speech and non-speech audio such as sound effects and speaker identification, while shortshort captions carry spoken words only. W3C also says automatic captions only meet accessibility requirements once confirmed fully accurate, so proofread them.

Will YouTube or TikTok add their own captions on top of mine?

Both platforms can generate automatic captions: YouTube publishes them on Shorts when available, and TikTok generates them for uploaded videos. Check the post after upload and edit or remove the platform captions if the text appears twice.

Can I upload shortshort captions as a separate caption track?

Yes. 'Export captions' saves an SRT file built from the same pages as the burned-in captions and timed on the short's timeline. YouTube accepts basic SRT files without style markup, so the visual style does not carry over.

Can shortshort translate captions?

No. Captions stay in the spoken language, and there is no translation or dubbing. Whisper supports 98 languages according to OpenAI, but accuracy varies by language, so check the word list before rendering.

Sources

The official pages this guide relies on, with the day each one was read.

Related

All 5 guides

Cut your own long video into shorts.

60 credits are free at sign-up, no card. One credit is one minute of source video, whatever the number of shorts.