The 9:16 vertical video guide: safe zones, hooks, and framing for phones
Vertical video is the 9:16 aspect ratio — 1080×1920 pixels at standard resolution — designed to fill a phone screen held the way people actually hold phones. Making it well comes down to four crafts: keeping critical content inside the safe zones the app interface does not cover, earning attention in the first one to two seconds, framing for a screen the size of a hand, and exporting the right dimensions for each surface. This guide covers all four, plus the question that decides quality before you shoot or generate a single frame: whether to create in 9:16 natively or crop a 16:9 original.
Why 9:16 is its own craft, not rotated 16:9
A vertical frame is not a horizontal frame turned sideways; it is a different compositional problem. A 16:9 frame is wide enough to hold a subject and its context — a person and the room, a product and the scene. A 9:16 frame has roughly a third of that horizontal room, so it holds one thing at a time: one face, one product, one line of text. Everything in vertical craft follows from that constraint.
The viewing situation is different too. Vertical video plays full-screen by default, sound often starts muted, and the viewer's thumb is already resting on the screen, one flick away from the next video. You are not competing with other channels; you are competing with the swipe. That is why hooks, captions, and framing carry more weight here than in any horizontal format.
Safe zones: the app's interface eats your frame
Every short-video feed draws its own interface on top of your video: a caption block and account name along the bottom, engagement buttons stacked on the right edge, status and navigation elements at the top. Your video renders underneath all of it. Anything you place in those regions — a subtitle, a logo, the product you are pointing at — gets covered for the viewer even though it looked perfect in your editor.
The exact overlay differs by app and changes with redesigns, so the durable habit is a margin discipline rather than a memorized pixel map:
- Bottom ~20% of the frame: treat as unusable for text. This is where captions, usernames, and audio attribution sit in most feeds.
- Right ~15%: keep faces, products, and text clear of it. The engagement button stack lives here.
- Top ~10%: avoid text; leave breathing room for status bars and search elements.
- The middle 60–70%, biased slightly above center: this is your real canvas. Faces, headlines, and the thing you are selling live here.
On a 1080×1920 frame, that means keeping essential content roughly inside a centered column about 900 pixels wide and away from the bottom 380 pixels. Our social media sizes reference keeps the current dimensions and safe-area notes per surface.
The practical workflow: build a safe-zone habit into your title placement. In PrismPoster's timeline editor, place text in the upper-middle of the frame by default and only move it after checking what it competes with.
The hook: the first two seconds decide everything
Vertical feeds are swipe-driven, and viewers decide whether to stay almost immediately. The opening of a 9:16 video has one job: give a reason to not swipe before the viewer's thumb finishes its resting motion. Some patterns that reliably do this:
- Start mid-action. Open on the most visually interesting moment, not the setup. If the payoff is a transformation, show a flash of the end state first.
- Put the promise on screen as text. Since sound often starts muted, a short on-screen line — the question the video answers — works even in silence. Keep it inside the safe zone and readable at phone size.
- Cut early and visibly. A camera or scene change inside the first second signals that the video moves.
- Never open with a logo card. A branded intro is a request to leave. Identity belongs at the end or in a corner, not in the two seconds that decide retention.
Because AI video clips run short — PrismPoster's Video Studio generates 4 to 10 seconds per clip, and single AI generations stay short everywhere — the hook discipline maps naturally onto generated footage: make your first clip the hook clip, generated specifically to be arresting, and assemble the rest behind it on the timeline.
Framing for a screen the size of a hand
Compositional rules that work on a monitor fail on a phone. Adjustments that matter:
One subject per frame. The vertical frame cannot hold a two-shot comfortably. If two people talk, cut between singles rather than cramming both into frame.
Frame tighter than feels natural. A "medium shot" on a phone reads like a wide shot. Faces should be large — head-and-shoulders framing, eyes in the upper third. Detail shots (hands, products, textures) can fill the entire frame.
Compose in vertical thirds. Think of the frame as top, middle, bottom bands: subject in the middle band, headline text in the top band, and the bottom band left for the feed's own interface. When prompting generated video, say so explicitly — "vertical composition, subject centered, headroom above" steers the model better than hoping a horizontal-minded prompt crops well.
Make text huge. Type that looks tastefully sized in your editor is illegible at arm's length on a 6-inch screen. If a caption feels slightly too big on your desktop preview, it is probably right. If you rely on spoken content, burn in subtitles — AI captioning makes this a minutes-long job, and muted autoplay makes it mandatory.
Respect vertical motion. Horizontal pans waste the axis you do not have. Vertical moves — tilt-ups, rises, falling objects, top-to-bottom reveals — use the frame's actual shape.
Sizes per surface
The good news: 9:16 at 1080×1920 is the universal answer for short vertical video. TikTok videos, Instagram Reels, YouTube Shorts, Facebook Reels, and Snapchat all take it. Vertical variation across surfaces today is mostly about length limits and safe-area details rather than dimensions — the social media sizes page tracks the current specifics per surface.
Two ratios get confused with 9:16 and deserve a warning:
- 4:5 (1080×1350) is the portrait feed format on Instagram — taller than square, shorter than full vertical. A 9:16 video shown in a 4:5 feed slot gets its top and bottom trimmed — exactly where your text is if you ignored the safe zones.
- 1:1 (1080×1080) square still has uses in feeds and ads but is not a vertical format; treating it as interchangeable with 9:16 loses a third of your canvas.
If you are converting between any of these — checking what a 16:9 master becomes at 4:5, or what resolution a crop leaves you — the aspect ratio calculator does the arithmetic, including the resulting pixel dimensions after a crop.
On resolution: 1080×1920 is the standard delivery size, and most feeds compress anything larger back down. Generating or editing at higher resolution still helps when you plan to crop, zoom, or punch in during the edit — our resolution guide covers when the extra pixels pay for themselves and when they are wasted credits.
Native 9:16 versus cropping 16:9
The most consequential vertical-video decision happens before any footage exists: do you create in 9:16, or create in 16:9 and crop?
Cropping costs you two-thirds of the frame. A center crop of a 16:9 image to 9:16 keeps about a third of the horizontal pixels. Everything composed for the wide frame — the second person, the context, the negative space — is gone, and off-center subjects end up half-out of the vertical frame. You can track the crop shot by shot in an editor, but you are salvaging, not composing.
Cropping costs resolution. Cut a 1920×1080 frame down to a 9:16 slice and you keep a 608×1080 region, which then has to stretch to fill a 1080×1920 screen. That is an upscale of what survived, and it shows.
Native 9:16 sidesteps both problems. When the frame is vertical from the first pixel, the model (or camera) composes for it: subject placement, headroom, and motion all assume the tall frame. PrismPoster's Video Studio generates 9:16 natively — text-to-video and image-to-video both, at 360p through 4K — so a vertical clip is composed vertical, not rescued from a horizontal one. The image-to-video path is particularly clean for this: generate a 9:16 still in the Image Studio, get the composition exactly right while iterating cheaply on images, then animate the keeper — the workflow our image-to-video guide walks through. (One honest limit: image-to-video declines raw photos of identifiable real people; generated stills, products, art, and landscapes animate without restriction.)
When does cropping make sense? When the 16:9 already exists and reshooting or regenerating is off the table — repurposing a landscape master into a vertical cut. Then the job is damage control: track the subject through the crop, re-place all text inside the vertical safe zones, and accept that some shots will need to be replaced rather than cropped. If you are producing both orientations of a planned piece, generate each natively instead of cropping one from the other; two compositions beat one compromise.
A vertical workflow that holds together
Pulling the craft into one sequence you can actually run:
- Write the hook first — the on-screen line and the opening visual — before anything else.
- Generate or shoot in 9:16 natively; iterate composition on cheap stills before spending on video generation.
- Assemble on the timeline with the hook clip first; keep every cut earning the next second.
- Add captions, sized for a phone, placed in the upper-middle safe zone.
- Export at 1080×1920 and check the frame against the safe-area notes for wherever you plan to publish it.
For the ideas layer on top of this mechanical layer — formats and concepts that suit short vertical video — the Reels-style generation guide picks up where this one stops.
Frequently Asked Questions
What size is a 9:16 vertical video?
1080×1920 pixels is the standard. The ratio means 9 units wide for every 16 tall — a 16:9 frame rotated in proportion, though not in composition. All the major short-video feeds accept 1080×1920, and most compress larger uploads down to something near it.
What are the safe zones for vertical video?
Keep essential content out of roughly the bottom 20%, right 15%, and top 10% of the frame, because feed interfaces draw captions, buttons, and navigation there. The reliable canvas is the middle of the frame, biased slightly above center. Exact overlays vary by app and change over time, so check current dimensions before finalizing text placement.
Is it better to shoot vertical or crop from 16:9?
Native vertical, whenever you control the source. Cropping 16:9 to 9:16 discards about two-thirds of the horizontal frame and most of the resolution, and the composition was never designed for the tall shape. Crop only when the horizontal footage already exists and cannot be replaced.
How long should a vertical video be?
As long as it stays interesting and not a second longer — for most short-feed content that means well under a minute, and the first two seconds matter more than the total length. AI generation naturally produces short clips of a few seconds each, which you then sequence on a timeline into whatever runtime the idea deserves.
Can AI generate vertical video directly?
Yes. PrismPoster's Video Studio generates in 9:16 natively from a text prompt or a starting image, at 360p up to 4K, in 4-to-10-second clips. Because the model composes for the vertical frame, the result needs no cropping — subject placement and headroom are already right for a phone screen.
Try it on a real vertical project
The fastest way to internalize safe zones and hook timing is to finish one vertical video and watch it on your own phone. Check your dimensions against the social media sizes reference, run any crop math through the aspect ratio calculator — both free, no account needed. When you want to generate rather than calculate, a free account comes with a one-time grant of 200 credits, no card required: enough to iterate a 9:16 composition on stills, animate the keepers into vertical clips, and cut them together with captions on the timeline today. The credits never expire, so the project can wait for the idea.
Turn ideas into high-definition AI video clips & timeline cuts
Generate text-to-video and image-to-video in 360p up to 4K, extend scenes per second, and assemble complete multi-track videos with auto-captions and zero desktop install.