Captions are timed text that represents spoken dialogue and may also identify speakers or describe meaningful sounds. Subtitles usually represent or translate spoken dialogue for viewers who can hear the audio. That is the textbook distinction, and it comes with an honest footnote: it is not universal. Some platforms and regions use the terms interchangeably, and same-language subtitles do the same job as captions.
Which leads to the actually useful rule: do not rely on the label. Define what the text must contain, which language it uses, how it is timed, and whether viewers can turn it off.
The distinction, when it holds
Captions assume the viewer may not hear the audio, so they can carry more than words: "[door closes]", "[upbeat music]", "MAYA: The project is ready for review." Subtitles assume the audio is heard and the words need representing or translating. Whether every meaningful sound gets described is a specification decision, not something the word "captions" guarantees.
Subtitles are not always translated, either. English speech with English subtitles is intralingual and common; English to French is interlanguage. Both are subtitles.
Open versus closed is about delivery, not content
Open captions are burned into the image: always visible, identical everywhere, no player support needed, and impossible to turn off or correct after export. Closed captions ship as a separate selectable track: viewer-controlled, correctable, machine-readable, and dependent on the player rendering them properly.
The key insight: open and closed describe the delivery method, not the content type. Burned-in subtitles exist; selectable captions exist; every combination is real. For short-form feeds, burned-in text is the norm because platform players and autoplay make selectable tracks unreliable; for long-form and multilingual work, tracks win.
The neighbours
SDH commonly means subtitles for the deaf and hard of hearing: subtitle files that also carry caption conventions like sound descriptions and speaker labels. A transcript is the written record without timed segments; it can become captions, but it is not captions until it has been corrected, segmented, timed and formatted. On-screen text is the broadest bucket: headlines, labels, "Workflow updated, March 2026." None of it is captioning; all of it shares the frame with your captions.
File formats prove nothing here: SRT and WebVTT are containers. Do not identify the deliverable by the filename.
How to request the right thing
Instead of ordering "captions" or "subtitles", specify: the language; same-language or translation; speaker labels yes or no; sound descriptions yes or no; burned-in or selectable; the file format if separate; the styling; who reviews; and the destination. Nine lines, and every future misunderstanding is gone. The final deliverable may well be more than one output: a burned-in social version plus a clean master plus a subtitle file.
Where ReadyForm fits
ReadyForm sits firmly on one side of this map: same-language, word-timed captions rendered into the video, in the style your brand kit defines, corrected through the story view where the transcript and the captions are the same thing. It does not translate and does not export separate subtitle files, so translated or selectable deliverables remain a separate step in your workflow. See how captions work in the product.