Sparround

Captions and on-screen text: designing for the muted viewer

Most people watch short video with the sound off — on transport, at work, at night next to someone sleeping. That imposes a design requirement: the meaning must land fully with no audio.

[PRACTICE] The test is simple: after the edit, watch the video muted on a phone. If you can follow what is happening, the design works. If you cannot, the video is empty for much of your audience.

There are two separate elements to this, and they must not be conflated: captions (the text of what is said) and on-screen text (emphasis, the hook, step numbers).

ParameterWhat worksWhat does not
FontBold, simple, sans-serifDecorative, script or thin fonts
SizeReadable at arm's length on a phoneA size that looks "fine" on a desktop
ContrastWhite text with a black outline, or a black box behind itWhite text alone — it disappears on light backgrounds
PositionAbove the middle lineAt the very bottom — buried under the caption and buttons
Line countOne or two lines, in short chunksFour- or five-line blocks — nobody reads them
AnimationNone, or the simplest availableEvery word flying in separately — it slows reading
LanguageOne language, consistentlyHalf in one language, half in another — it hurts both

Do not trust auto-captions — correct them.

[OFFICIAL] CapCut generates captions automatically, and its own documentation acknowledges the output is not always accurate. In practice the errors cluster around proper nouns, technical terms, sentences mixing two languages and fast speech.

The problem is sharper in Azerbaijani. The practical approach: generate, then always read and fix — especially the key terms, because they matter both to the viewer and to search.

A key term left misspelled is both an unreadable caption and lost search traffic.

Captions and on-screen text do different jobs — separate them.

[PRACTICE] The most common design failure: the creator generates auto-captions and adds nothing else. The result has no hook on screen, no visible steps, no emphasis — just a stream of transcribed speech.

Separate the two like this:

  • Captions — small, above the middle line, synced to speech, running throughout
  • On-screen text — large, in the upper third, used sparingly, only at the moments that matter: the hook (0-2s), step numbers ("1/3"), a key figure or the result

Their size and position must differ — otherwise the screen turns to noise and neither gets read.

Pick one style and stay with it. [PRACTICE] Using a different caption style on every video makes the account harder to recognise. Same font, same colour, same position — and a viewer scrolling the feed knows the video is yours before reading a word.

This is one of the main payoffs of the master project template from earlier in this stage: the style is chosen once and repeats itself.

🛠 Practice task

Watch your latest video muted on a phone and run the six-question test. Fix everything that comes back "no".

Then finalise the caption style in your master template: bold sans-serif, an outline or background box, positioned above the middle line. Add a separate hook text box — 1.5-2× the caption size, in the upper third.

It is done when all six test questions answer "yes" and the template holds two distinct text elements (captions and hook) at different sizes and positions.

📚 Sources and documentation