ppl.studio

Product photography & image generation

Photographing and generating product imagery — formats, scene construction, marketplace requirements, and the models now producing it.

48 terms in this topic.

  • AI background generation

    AI background generation replaces or creates the environment behind a subject while preserving the subject itself — putting a product cut from a plain studio shot into a kitchen, a shelf, or an outdoor scene.

  • Amazon Brand Storefront

    An Amazon Brand Storefront (also called Amazon Store) is a free, customizable multi-page brand destination within Amazon that is available to brand-registered sellers and vendors. Brand Storefronts function as a brand's homepage on Amazon—a curated shopping experience where customers can browse the full product catalog, learn the brand story, and discover products by category without competitor ads appearing on the page.

  • Amazon product photography

    Amazon product photography is imagery produced to Amazon's listing requirements and to the behaviour of Amazon shoppers, which differ enough from a brand's own site to warrant separate treatment. The main image is tightly constrained: a pure white background, the product occupying roughly 85 percent of the frame, no props, text, watermarks, or logos that are not on the product itself.

  • Beauty photography

    Beauty photography is a specialized genre focused on cosmetics, skincare, haircare, and personal care products. It encompasses close-up product shots, shade swatches on skin, application tutorials, before-and-after imagery, and lifestyle shots of people using beauty products. Beauty photography demands model diversity (skin tones, ages, features) and seasonal adaptability for collection launches, making it one of the most content-intensive categories in e-commerce.

  • Brand photography

    Brand photography is the coherent body of imagery representing a brand across its site, campaigns, social channels, and sales material — as distinct from product photography, which documents individual items for a listing. Where product photography answers what this is, brand photography answers who this is for and what it feels like to own.

  • Consistent character

    Consistent character refers to a generative AI workflow that produces the same recognizable person—same face, same body proportions, same wardrobe baseline—across dozens or hundreds of generated images and video clips.

  • ControlNet

    ControlNet is an open-source neural network architecture (introduced in 2023) that conditions an image-generation model on a structural reference — a pose skeleton, depth map, edge map, segmentation map, or sketch — alongside the text prompt. It is the technical foundation for 'I want this exact pose with this product on this background' control in modern AI image generation.

  • Diffusion model

    A diffusion model is a class of generative AI that creates images (or video, audio, 3D) by starting from pure random noise and iteratively 'denoising' it into a coherent output, guided by a text prompt or a reference image.

  • E‑commerce photography

    E-commerce photography is the specialized discipline of creating product imagery optimized for online selling across websites, marketplaces, and advertising platforms. It encompasses multiple image types, each serving a distinct purpose in the buyer journey: primary images (clean, white-background shots required by Amazon and Google Shopping as the main listing image), lifestyle and in-use photos (showing products in real-world contexts with people to drive emotional connection), detail and close-up shots (highlighting materials, textures, and craftsmanship), scale and comparison images (showing product size relative to familiar objects), infographic images (overlaying feature callouts, dimensions, and key benefits), and user-generated content (authentic customer or creator photos that serve as visual social proof).

  • Flat-lay photography

    Flat-lay photography is a style where objects are arranged on a flat surface and photographed from directly above. It's widely used in e-commerce, beauty, food, and lifestyle marketing to showcase products in curated, aesthetically pleasing compositions. Flat-lays are popular on Instagram, Pinterest, and product pages because they allow brands to show multiple products, ingredients, or accessories in a single, shareable image.

  • Flux 2

    Flux 2 is Black Forest Labs' next-generation image model family, succeeding the original Flux Pro / Flux Dev / Flux Schnell lineage that defined 2024 open-weights image generation. Flux 2 keeps the open-weight strategy that made Flux 1 the default backbone for indie product-photo tools and self-hosted creative pipelines, while shipping major improvements in photorealism, identity-lock under LoRA fine-tuning, multi-subject scenes, and text rendering.

  • Food photography

    Food photography is a specialized genre of commercial photography focused on making food and beverage products look appetizing and aspirational. It typically involves food stylists, specialized lighting, and careful composition. In e-commerce and marketing, food photography is used for product pages, social media, menus, packaging, and advertising.

  • Generative fill

    Generative fill is the user-facing name for the unified inpaint-or-extend operation introduced by Adobe in Photoshop (May 2023, on Adobe Firefly) and now widely copied by every major AI image editor. The user selects an area — either inside the image (inpaint) or outside the current canvas (outpaint / expand) — types a prompt or leaves it blank, and the model regenerates that region.

  • Ghost mannequin

    Ghost mannequin (also called the invisible or hollow-man effect) is a product-photography technique for apparel in which the garment is shot on a mannequin and the mannequin is then edited out, leaving the clothing holding its own three-dimensional shape as if worn by an invisible body.

  • GPT Image

    GPT Image is OpenAI's native image-generation model that ships inside ChatGPT and the OpenAI API as the successor to DALL-E. Where DALL-E was a standalone diffusion model called by an external tool, GPT Image is multimodal-native: the same model that handles text reasoning also generates the image, which yields dramatically better text-rendering, prompt fidelity, and conversational editing.

  • Hero image

    A hero image is the single most prominent image on a page — the first visual in a product gallery, the banner at the top of a landing page, or the thumbnail representing a listing in search results. It carries disproportionate weight because it is processed before any text and often decides whether the rest of the page is read at all.

  • Image generation

    Image generation is the process of creating new images from scratch using AI models—typically diffusion models like Stable Diffusion, DALL-E, Midjourney, or Flux—based on text prompts, reference images, or a combination of both. Modern image generation models can produce photorealistic scenes, artistic illustrations, product visualizations, and stylized graphics at resolutions suitable for professional marketing use.

  • Image prompt engineering

    Image prompt engineering is the discipline of writing text prompts that reliably produce the intended visual output from a generative image model. It is distinct from LLM prompt engineering because image models respond to a different vocabulary: subject + setting + camera type + lens + lighting + composition + style — in that order, weighted from most to least important.

  • Image-to-image (img2img)

    Image-to-image (img2img) is the AI generation mode where a source image is used as the starting point for a new generation, rather than starting from random noise. It is the core technique behind 'restyle this photo,' 'put this product in a different scene,' and 'turn this sketch into a polished render.' The strength of the source image's influence is controlled by a denoising parameter — low denoising keeps the original layout and details, high denoising lets the model reinvent more freely. img2img is the technical mechanism for product-placement AI photography: the product photo is fed in as the source, the model preserves its silhouette and material properties, and the surrounding scene is generated fresh.

  • Imagen 4

    Imagen 4 is Google DeepMind's flagship text-to-image model, the successor to Imagen 3, launched at Google I/O 2025. It is available via the Gemini API, Vertex AI, and inside consumer Google products (Gemini app, ImageFX, Google Slides).

  • Inpainting

    Inpainting is the AI image-editing operation that regenerates a masked region of an existing image based on a text prompt, while leaving the unmasked region pixel-identical. The underlying mechanism is diffusion conditioned on both the surrounding pixels and the prompt, so the regenerated region blends seamlessly into the original.

  • Lifestyle content

    Lifestyle content is marketing content that shows products or services being used in real-world, aspirational settings rather than isolated product shots. By placing products in context—being worn, used, or enjoyed by relatable people in everyday or aspirational scenarios—lifestyle content helps consumers envision themselves using the product.

  • Lifestyle photography

    Lifestyle photography shows products or people in real-world, contextual settings that help viewers imagine the product in their own lives—someone applying skincare in a sunlit bathroom, a runner lacing up shoes on a forest trail, or a family gathered around a dinner table with a new cookware set.

  • Lookbook

    A lookbook is a curated collection of styled photographs that showcase a brand's products—typically fashion, beauty, home, or lifestyle items—in aspirational, editorial-quality settings. Unlike a product catalog that focuses on specifications and pricing, a lookbook tells a visual story: it shows how products look when worn, styled, or placed in real-world contexts, communicating the brand's aesthetic and the lifestyle it represents.

  • Marketplace photography

    Marketplace photography refers to product imagery optimized for online marketplace platforms like Amazon, Walmart Marketplace, eBay, Etsy, and Target Plus. Each marketplace has specific image requirements (dimensions, background rules, image count limits) and buyer expectations that differ from direct-to-consumer or social media photography.

  • Midjourney

    Midjourney is one of the most widely-used commercial text-to-image AI platforms, known for its distinctive painterly aesthetic, strong prompt adherence on artistic concepts, and a large active community of prompt-engineers. The platform launched in 2022 and progressed through V4, V5, V6, and V7 model generations, each iteration improving photorealism, in-image text rendering, and identity-lock under reference images.

  • Multimodal AI

    Multimodal AI is a model architecture that processes and generates across multiple input/output modalities — text, image, video, audio — in a single unified system. The major multimodal models as of 2025: Google Gemini 2.5 Flash/Pro (text + image + video + audio), OpenAI GPT-4o and GPT-5 (text + image + voice + video), Anthropic Claude (text + image), Meta Llama 3.2 Vision (text + image).

  • Nano Banana (Gemini image model)

    Nano Banana is the community nickname for Google's Gemini 2.5 image generation and editing model, released in 2025 and widely adopted for product-photo workflows because of its unusually strong identity consistency, prompt adherence, and ability to faithfully render real products from reference images.

  • Negative prompt

    A negative prompt is text that tells an image-generation model what to avoid producing. It is supplied alongside the main (positive) prompt and acts as a guidance signal pulling the generation away from listed concepts. Common negative-prompt patterns: 'extra fingers, distorted face, watermark, logo, low quality, blurry' for photorealistic UGC; 'cartoon, anime, illustration' for forcing photoreal style; 'busy background, clutter' for clean compositions.

  • Outpainting

    Outpainting is the AI image-editing operation that extends an image beyond its original canvas — generating new pixels above, below, or to the sides of the existing image, conditioned on what is already there plus a prompt. It is the inverse of inpainting: instead of regenerating a region inside the image, it generates new regions outside it.

  • Product detail shot

    A product detail shot is a close-up photograph highlighting specific features, textures, materials, or craftsmanship of a product. Detail shots serve as supporting images on product pages, complementing hero shots and lifestyle images by giving shoppers a closer look at quality indicators like stitching, finishes, hardware, or material grain.

  • Product hero shot

    A product hero shot is the primary, attention-grabbing image of a product used as the main visual in listings, advertisements, landing pages, and marketing materials. It is the single most important product image because it creates the first impression and determines whether a shopper clicks, scrolls past, or engages.

  • Product listing photos

    Product listing photos are the images used on marketplace and e-commerce product detail pages to showcase a product to potential buyers. On Amazon, Shopify, Walmart Marketplace, Etsy, and other platforms, product listing photos are the single most influential factor in purchase decisions—shoppers cannot physically examine products, so the images must communicate quality, scale, features, and use cases.

  • Product mockup

    A product mockup is a realistic visual representation of a product placed in a context or scene for marketing purposes. Mockups show how a product looks in real-world settings—on a shelf, in someone's hand, on a desk, or in a lifestyle environment—without requiring physical photography.

  • Product photography

    Product photography is the creation of images that showcase a product for e-commerce listings, advertisements, catalogs, social media, and marketing collateral. It encompasses several distinct styles: white-background studio shots (required by Amazon and Google Shopping as primary images), lifestyle photography (products shown in real-world contexts), flat-lay compositions (products arranged on surfaces and shot from above), detail and macro shots (close-ups highlighting textures, materials, and craftsmanship), and product-in-use imagery (people actively using or wearing the product).

  • Product staging

    Product staging is the process of arranging a product in a styled environment for photography—setting up backgrounds, props, lighting, and complementary objects to create a composed scene. In traditional photography, staging is one of the most time-consuming and expensive steps, requiring physical spaces, stylists, and props.

  • Product visualization

    Product visualization is the process of creating images, 3D renders, or composited scenes that show a product in context before or without conducting physical photography. It is widely used in e-commerce, advertising, and product development to present items in lifestyle settings, on model figures, or in branded environments.

  • Product-in-scene

    Product-in-scene photography shows a product placed in a realistic, contextual environment rather than isolated on a plain background — a skincare bottle on a bathroom counter, a laptop on a cafe table, a supplement tub beside a gym bag. It converts better than white-background imagery for two specific reasons.

  • Prompt engineering

    Prompt engineering is the practice of writing inputs to AI models in ways that reliably produce desired outputs. The discipline emerged from research on large language models and has since expanded to image generation (image-prompt-engineering), video generation, agentic AI, and multimodal systems.

  • Scene generation

    Scene generation is the AI-powered creation of contextual backgrounds, environments, and settings for products and people in marketing imagery. Rather than shooting on location or building physical sets, brands use scene generation to place products in kitchens, bathrooms, offices, outdoor settings, or any environment that matches their target audience.

  • Scene prompt

    A scene prompt is the natural-language description that tells an image generation model what scene to render—the location, lighting, mood, time of day, camera framing, and subject pose. Scene prompts are distinct from style prompts (which control the look: 'cinematic,' 'film grain,' 'soft pastel') and identity references (which control who or what appears: a product photo, a face).

  • Stable Diffusion

    Stable Diffusion is the most widely-deployed family of open-weights text-to-image diffusion models, originally released by Stability AI and Runway in August 2022. The release was a defining moment for AI image generation because it shipped under a permissive license with downloadable model weights — enabling the entire ecosystem of self-hosted image pipelines, LoRA fine-tuning, ControlNet conditioning, IP-Adapter face-locking, and the commercial tools built on top (Leonardo, Playground, fal.ai, Replicate, NightCafe, countless niche product-photo tools).

  • Style transfer

    Style transfer is the operation of taking the visual style of one image (color palette, lighting, brush stroke, composition logic) and applying it to the content of another. The classical 2015 implementation (Gatys et al., neural style transfer with VGG-19) gave us Prisma-style filter effects; the diffusion-era implementation (IP-Adapter, ControlNet Reference, StyleAlign, Flux Style transfer) is far more powerful and content-preserving.

  • Text-to-image

    Text-to-image is the AI workflow of producing an image directly from a written description, with no reference image required. It is the foundational capability of modern image-generation models from Midjourney to DALL·E 3 to Flux to Imagen.

  • Virtual photoshoot

    A virtual photoshoot is a product or lifestyle photography session conducted entirely through AI and digital tools, without a physical studio, photographer, models, or physical staging. Virtual photoshoots use AI image generation, product visualization, and persona consistency technology to produce images that are visually comparable to traditional studio photography—lifestyle product shots, model-with-product compositions, scene-based product imagery, and brand storytelling visuals.

  • Visual preset

    A visual preset is a saved bundle of generation parameters—lighting, camera angle, location category, color grading, mood, composition rules—that can be applied to any new prompt with a single click, ensuring visual consistency across a batch of generated images.

  • Visual storytelling

    Visual storytelling is the practice of conveying a narrative, message, or brand story primarily through images, video, and visual sequences rather than text. In marketing, visual storytelling uses lifestyle photos, product-in-use imagery, and sequenced visuals (carousels, storyboards) to communicate how a product fits into someone's life.

  • White-background product photography

    White-background product photography is the shooting and editing of a product against a pure white field (commonly the pure-white value RGB 255,255,255) so the item appears cleanly isolated with no distracting setting — the format marketplaces require for main catalog images and the default for comparison-shopping grids.

Back to the full glossary