Captions and on-screen text: designing for the muted viewer
Most people watch short video with the sound off — on transport, at work, at night next to someone sleeping. That imposes a design requirement: the meaning must land fully with no audio.
[PRACTICE] The test is simple: after the edit, watch the video muted on a phone. If you can follow what is happening, the design works. If you cannot, the video is empty for much of your audience.
There are two separate elements to this, and they must not be conflated: captions (the text of what is said) and on-screen text (emphasis, the hook, step numbers).
| Parameter | What works | What does not |
|---|---|---|
| Font | Bold, simple, sans-serif | Decorative, script or thin fonts |
| Size | Readable at arm's length on a phone | A size that looks "fine" on a desktop |
| Contrast | White text with a black outline, or a black box behind it | White text alone — it disappears on light backgrounds |
| Position | Above the middle line | At the very bottom — buried under the caption and buttons |
| Line count | One or two lines, in short chunks | Four- or five-line blocks — nobody reads them |
| Animation | None, or the simplest available | Every word flying in separately — it slows reading |
| Language | One language, consistently | Half in one language, half in another — it hurts both |
Do not trust auto-captions — correct them.
[OFFICIAL] CapCut generates captions automatically, and its own documentation acknowledges the output is not always accurate. In practice the errors cluster around proper nouns, technical terms, sentences mixing two languages and fast speech.
The problem is sharper in Azerbaijani. The practical approach: generate, then always read and fix — especially the key terms, because they matter both to the viewer and to search.
A key term left misspelled is both an unreadable caption and lost search traffic.
Captions and on-screen text do different jobs — separate them.
[PRACTICE] The most common design failure: the creator generates auto-captions and adds nothing else. The result has no hook on screen, no visible steps, no emphasis — just a stream of transcribed speech.
Separate the two like this:
- Captions — small, above the middle line, synced to speech, running throughout
- On-screen text — large, in the upper third, used sparingly, only at the moments that matter: the hook (0-2s), step numbers ("1/3"), a key figure or the result
Their size and position must differ — otherwise the screen turns to noise and neither gets read.
Pick one style and stay with it. [PRACTICE] Using a different caption style on every video makes the account harder to recognise. Same font, same colour, same position — and a viewer scrolling the feed knows the video is yours before reading a word.
This is one of the main payoffs of the master project template from earlier in this stage: the style is chosen once and repeats itself.
🛠 Practice task
Watch your latest video muted on a phone and run the six-question test. Fix everything that comes back "no".
Then finalise the caption style in your master template: bold sans-serif, an outline or background box, positioned above the middle line. Add a separate hook text box — 1.5-2× the caption size, in the upper third.
It is done when all six test questions answer "yes" and the template holds two distinct text elements (captions and hook) at different sizes and positions.
📚 Sources and documentation
- Fixing inaccurate auto-captions — CapCutofficialcapcut.com
Why caption errors occur and how to correct them.
- How to add subtitles in CapCutofficialcapcut.com
Generating captions, styling them and changing their position.
- Creative best practices for performance adsofficialads.tiktok.com
The safe zone and where text is supposed to sit.