Vertical video is any frame taller than it is wide, and it is a category, not one size. The family: 9:16 (1080 by 1920) for tall full-screen video, 4:5 (1080 by 1350) for portrait feeds, 3:4 (1080 by 1440) in between, and 2:3 (1080 by 1620) for photography-led layouts. There is no single universal vertical-video size.
The orientation test is arithmetic: height greater than width is vertical, equal is square, less is horizontal. Which settles a common one: 1080 by 1080 is square, not vertical. And "portrait" is the same thing as vertical in practice; the ratio and resolution are the details that matter.
The four shapes, in use
9:16 owns the full-screen feed: maximum vertical coverage, room for a speaker plus captions, and the least width in the family, which is exactly its challenge with two-person scenes, wide products and desktop interfaces. 4:5 trades height for width: better for broader gestures, side-by-side elements and feed placements, and it does not fill a 9:16 surface without cropping or a designed background. 3:4 sits between them, less common as a short-form master; check the destination before adopting it. 2:3 leans photographic: taller than 4:5, wider than 9:16, and not every destination presents it identically.
Choosing between them is five questions: where will it be published, what must stay visible, how was the source recorded, will several formats be needed, and how much text sits on screen.
One master, several verticals
The efficient workflow is one approved content master and separate format versions, sharing the takes, message, audio and brand assets, and differing in everything spatial: crops, caption groups and positions, title layouts, CTA placement, scale. One source does not mean one automatic export; each version deserves its own scene walk, because the important subject shifts per scene and a single global crop always guillotines something.
The hard cases repeat across formats. Multiple speakers: alternate close-ups, a stacked layout, active-speaker reframing, or keep it horizontal; automatic tracking gets checked, always. Products: a crop that keeps the person but loses the demonstrated object has failed the scene's purpose. Screen recordings: crop, zoom, one step at a time, or a diagram; a shrunken desktop serves nobody.
And captions travel worst of all: line widths, breaks and safe positions differ per ratio, so every version gets its own caption pass rather than inherited pixel coordinates.
Source quality sets the ceiling
A 540 by 960 recording exported at 1080 by 1920 is the same limited detail wearing a bigger coat. Output dimensions do not recreate missing source information; record with enough resolution and working space, and the format versions stay sharp.
Where ReadyForm fits
ReadyForm asks for the aspect ratio at upload and composes the complete edit for it, captions included. Recording with a little room around the subject is what gives that composition its options. Every scene stays individually reviewable, which is where the per-scene reframing judgement above lands in practice. See how the edit is made.