ppl.studio

Video & short-form

Video formats and the generation techniques behind them — vertical, talking-head, lip-sync, and the specifications each surface expects.

19 terms in this topic.

  • Aspect ratio

    Aspect ratio is the proportional relationship between a media asset's width and height (e.g., 9:16, 1:1, 4:5, 16:9), and it's the single most important production decision in modern paid social — because every platform rewards content cropped natively for its dominant feed. 9:16 (vertical) is the standard for TikTok, Reels, Stories, YouTube Shorts; 1:1 (square) for Instagram Feed and many DTC product page formats; 4:5 (portrait) for Meta Feed at full attention; 16:9 (landscape) for YouTube and connected-TV.

  • B-roll

    B-roll is the supplementary footage cut in over the main shot — product close-ups, hands using the item, environment and detail shots, texture and packaging — used to illustrate what the primary audio is saying and to hide edits in the main take.

  • Hero video

    A hero video is the primary video asset featured prominently on a webpage, landing page, product page, or social media profile. Hero videos serve as the first visual experience visitors encounter and typically aim to communicate a brand's value proposition, demonstrate a product in action, or create an emotional connection within the first 3–5 seconds.

  • Hook rate

    Hook rate is the percentage of impressions that result in a viewer watching the first three seconds of a video ad, calculated as three-second video views divided by impressions. It isolates one job: whether the opening frame stops a scroll.

  • Image-to-video

    Image-to-video is a generative AI workflow where a still photo is the starting frame and the model produces a short video clip animating from that frame—either bringing the subject to life (a person speaking, an object moving, a camera panning) or holding the subject and animating only the environment (steam rising from coffee, leaves rustling, light shifting).

  • Instagram Reels

    Instagram Reels is Instagram's short-form vertical video format, distributed largely through algorithmic recommendation rather than to existing followers. That distribution model is the reason it matters commercially: a Reel reaches people who do not follow the account, which makes it the platform's main discovery surface, while feed posts and Stories skew toward an audience already acquired.

  • Kinetic captions

    Kinetic captions are animated, word-by-word subtitles synced to spoken audio, typically with bold typography, color highlights, and pop-in motion that emphasize each beat of the script. Popularized on TikTok and Reels and now standard across short-form video, kinetic captions improve both accessibility and watch-time — viewers who watch with sound off (estimated 60–85% of social video viewers) can still follow the message, and the animation itself functions as a visual hook that re-engages attention every few words.

  • Lip-sync

    Lip-sync (in AI video) is the process of mapping a generated or recorded audio track to the mouth movements of an on-screen subject so the visemes match the phonemes — i.e., the AI Expert's lips look like they're actually saying the words.

  • Shoppable video

    Shoppable video is video content that lets the viewer purchase a featured product directly from inside the video player — without clicking through to a separate product page. Major platforms shipped shoppable video formats between 2022 and 2025: TikTok Shop's in-feed video checkout, Instagram Reels shoppable tags, YouTube Shopping, Pinterest Shopping Pins on Idea Pins, and Amazon's Inspire feed.

  • Sora 2

    Sora 2 is OpenAI's second-generation text-to-video model, released in late 2025 as a major upgrade over the original Sora preview from early 2024. Sora 2 produces 1080p clips up to 60 seconds with native audio, supports image-to-video and video-to-video conditioning, and includes a 'storyboard' mode that lets users describe a sequence of beats rather than a single shot.

  • Storyboard

    A storyboard is the shot-by-shot plan for a video — each frame sketched or described with its visual, camera angle, on-screen text, and the line of voiceover or dialogue that runs under it — produced before anything is filmed or generated so the sequence is agreed while it is still cheap to change.

  • Talking-head video

    A talking-head video is a format in which a person speaks directly to camera, framed from roughly the chest up, usually holding or referencing a product. It is the dominant UGC ad format on TikTok, Reels, and Shorts because it mirrors how people already post on those platforms, and because direct eye contact carries a trust signal that product-only footage does not.

  • Text-to-video

    Text-to-video is the generative AI operation that takes a text prompt and produces a short video clip — typically 5–10 seconds, 720p–1080p, no input image required. It is the video analogue of text-to-image and the contrasting workflow to image-to-video (which conditions on a starting frame).

  • UGC video

    UGC video is user-generated video content—short-form clips, product demonstrations, testimonials, unboxing recordings, and lifestyle footage—created by consumers, creators, or AI tools rather than professional production teams. UGC video has become the dominant creative format in paid social advertising, with platforms like TikTok, Instagram Reels, and YouTube Shorts built entirely around short-form video consumption.

  • Veo 3 (and Veo 3.1)

    Veo 3 is Google DeepMind's text-and-image-to-video generation model, released in 2025 as the flagship of Google's video AI stack and used by Vertex AI, Gemini, and partner products like ppl.studio's Animate feature. Veo 3.1 is the late-2025 refresh that added synchronized audio, longer clip lengths, and tighter prompt adherence.

  • Vertical video

    Vertical video is video produced in a portrait aspect ratio (9:16, 4:5, or 2:3) for native consumption on mobile feeds where the device is held upright. The format went from niche to dominant between 2018 and 2025 with the rise of Snapchat Stories, Instagram Stories and Reels, TikTok, YouTube Shorts, and Pinterest Idea Pins — collectively the surface where Gen Z and Millennials now spend the majority of their content consumption time.

  • Video ad creative

    Video ad creative is the video asset itself — the footage, edit, audio, captions, and on-screen text that carry an advertising message, as distinct from the targeting and bidding that determine who sees it.

  • Video generation model

    A video generation model is a generative AI system that produces video output from a text prompt, an image, or a combination. The 2024–2025 generation includes Google Veo 3 and 3.1, OpenAI Sora and Sora 2, Runway Gen-3 and Gen-4, Luma Dream Machine, Pika 2, Kling, and Hailuo, plus open-source projects like LTX-Video and HunyuanVideo.

  • Voiceover

    A voiceover is recorded narration laid over visuals rather than spoken on camera — the voice carries the message while b-roll, product shots, or a muted talking-head take play underneath.

Back to the full glossary