B-roll in a talking-head video should appear when it helps explain, demonstrate, prove or contextualise what the speaker is saying, and the speaker should stay visible when facial expression, authority, emotion or direct delivery matters more than any additional visual. Do not cover every jump cut or every sentence; selective support beats constant coverage.
The premise the whole subject rests on: the speaker is part of the visual story. Facial expression, eye contact and body language communicate information no supporting clip can replace.
When to keep the speaker on screen
Six moments where cutting away costs more than it adds:
- Emotion: the story's heaviest sentence belongs to the face telling it.
- Authority: an expert claim lands harder from the expert.
- Trust: apologies, admissions and commitments need eyes.
- Nuance: qualified statements read differently with delivery attached.
- Personal connection: the moments that make the person the brand.
- Key transitions and the CTA: direct address closes better than a montage.
Do not hide the speaker during the strongest human moment only because B-roll is available.
When B-roll makes it better
The mirror list: show the product while it is being described, show the process while it is being explained ("recordings, selected takes, edit, review" as a visual beats hearing it twice), show the verified result when one is claimed, show the place when the place matters, and draw the abstract idea instead of describing it twice. And yes, a relevant clip can cover an edit; relevance first, coverage second.
The strong categories, in practice: product footage, focused screen recordings (crop or zoom around the relevant area, a full desktop is unreadable in a vertical frame), screenshots, diagrams, real process and location footage, and verified evidence. Generic stock and synthetic footage do not prove a specific result, and an unrelated stock location should never stand in for the real place.
The placement pattern that works
The reliable rhythm has three beats: the speaker introduces the point, the B-roll explains or demonstrates it, the speaker returns to interpret or conclude it. The visual starts once the complete idea has been spoken, and you come back to the face when the next thought needs delivery. Do not let B-roll run over unrelated narration just to avoid another cut.
How much in total? There is no universal percentage, frequency or maximum stretch of visible speaker. Every rule with a number in it ("every three seconds", "never longer than five") optimises the wrong thing. Judge B-roll by function, not by time.
Jump cuts, continuity and captions
A B-roll clip can hide the visible moment of a cut, and that is all it hides: inconsistent speech, a changed posture mid-sentence or a broken thought stay broken underneath. Alternatives to covering a cut: accept a clean jump cut, use a second angle, reframe, or find a better cut point.
Captions share the frame with everything else. Over a text-heavy screen recording, shorten the caption groups and check that neither layer covers the other; nobody reads two dense text layers at once.
By content type, briefly
Expert advice: visuals for the claims, face for the conclusions. Product explanation: recordings of the actual product, face for the why. Personal stories: real photos and places, emotion on the speaker. Founder content: the real office and the real prototype beat any stock scene. Testimonials: only genuine, approved material; never synthetic customers, never stock presented as the real person.
Where ReadyForm fits
ReadyForm makes the complete edit from talking-head takes and places supporting visuals with the pattern above in mind: found stock and your own Library media, timed to the sentence, with the speaker kept in view where the delivery carries the moment. Every placement is reviewable per scene, and removing one is one click. See talking-head editing with ReadyForm or how the edit is made.