To add captions online: upload or connect the footage, generate a transcript, correct the spoken text, divide it into readable units, adjust the timing and style, position the captions safely, and review the exported video from beginning to end.
Two rules do most of the work. Automatic captions are a starting point, needing checks on accuracy, timing, readability and placement. And AI can generate and time captions while you verify names, numbers, specialist terms and the complete meaning of each sentence.
Captions, subtitles, on-screen text
Captions represent spoken dialogue and relevant audio information. Subtitles represent or translate dialogue. On-screen text is everything else: titles, labels, statistics, CTAs. Not every visual text element is a caption, and the deliverable gets clearer the moment you say which one you mean.
The workflow
Upload and generate. Wait for processing to finish and caption the version you actually intend to keep; captions built from an incomplete edit get rebuilt.
Correct the text. The recurring failures: names, company and product names, numbers, prices, percentages, dates, currencies, abbreviations, URLs. The classic single-letter disaster: "Start your free trail today." And the boundary that matters most: do not rewrite the speaker silently. Captions must not make a materially different claim from the audio. When the spoken statement itself is wrong, correct the video or record a replacement rather than disguising it in text.
Group by meaning. Split at clauses and short sentences; never separate a first name from a last name, a number from its unit, or an article from its noun. Word counts vary with screen size, speaking pace, font size and position, so test rather than adopt a rule.
Fix the timing. Four failure modes: too early (it spoils the next point), too late (reading and hearing diverge), too fast (unreadable on a phone), too long (the video drags).
Style and position. Legible font over decorative, contrast tested on the hardest scene rather than the easiest, uppercase as an accent, highlights used sparingly enough that something still stands out. Position per scene, and always check captions at realistic mobile size.
Review in context, then export and inspect. Watch the whole video: the spoken "the first edit is complete, but it still needs to be reviewed" captioned as "the edit is complete" has changed responsibility and meaning. Then check the render for missing captions, shifted timing, broken lines, cropped text and fallback fonts.
Burned-in or separate
Burned-in captions guarantee consistent styling, always appear and never depend on the player, at the cost of being uncorrectable without a new export. Separate files can be toggled and edited, at the cost of styling limits and platform variation. Short-form feeds usually favour burned-in; multilingual and accessibility workflows often need files.
Where ReadyForm fits
ReadyForm makes word-timed captions with the edit itself, styled by your brand kit and rendered into the video. Corrections happen in the story view, where the transcript and the caption track are the same thing: fix the word, the caption follows, and the timing re-syncs. It does not translate or export separate subtitle files, so multilingual delivery stays a separate step. See the captions feature.