Video summary

Hice Anuncios Publicitarios con Gemini Omni por $0.90 Centavos (Así Quedaron)

Main summary

Key takeaways

Technology

Overview

This video tutorial walks through how to create low-cost, realistic UGC-style advertising videos using multiple AI tools—specifically OpenAI/GPT for images, Google Flow for video generation, ElevenLabs for consistent voice cloning, and CapCut for editing (including subtitle styling and ambience preservation). It also includes a subjective review of the resulting video quality.


Key Tech + Workflow (Step-by-Step)

1) Create an “influencer” image (OpenAI image generation / “image 2”)

The creator uses GPT chat with OpenAI image 2 to generate a vertical influencer image tailored for Reels/Stories.

  • The prompt is designed to resemble a real smartphone video frame:
    • casual and natural (not “professional production”)
  • Claimed output qualities:
    • natural facial features
    • realistic skin detail
    • phone-like interior car lighting

2) Generate a clean product render (remove condensation background)

Using GPT chat, the tutorial transforms a tonic can image (“Verona” brand; fictional) by:

  • removing condensation
  • switching to a pure white background
  • setting an aspect ratio described in the subtitles as roughly square-ish / 1:2.1

3) Generate the ad video with Google Flow (Google Omni / “UFO” model)

In Google Flow, the creator:

  1. Creates a new project
  2. Uploads both the influencer image and the product image
  3. Enables agent mode so the system can request changes and generate follow-up segments

Video generation settings (as described):

  • aspect ratio: 9:16
  • model: OmniFlash (per subtitles)

Script creation and segmentation:

  • A 24-second script is generated via GPT
  • It’s split into three ~8-second parts to guide scene generation

Important Flow feature: asset tagging

  • The creator highlights tagging assets so Flow knows which uploaded image to use for each action (using @ to attach:
    • the influencer image
    • the product can image)

Cost/credits (as shown/estimated):

  • first and second generations: 15 credits each (per approvals shown)
  • total estimate for the three clips: ~45 credits
  • creator estimate for real cost: ~90 cents or less, depending on the plan

4) Review of realism (quality analysis)

The creator reviews the generated footage and emphasizes:

  • natural gestures
  • realistic voice acting
  • organic facial expressions
  • lighting that matches a real phone-recorded scene

Caveat:

  • It’s not “100% perfect”—some AI artifacts remain
  • The creator expects realism to improve further with UFO and mentions something like “Sidans 2.0” (as transcribed)

5) Fix voice consistency (ElevenLabs / Seven Labs)

Because Flow’s voice consistency can be unreliable, the creator uses ElevenLabs to clone and enforce a consistent voice across all parts.

Steps described:

  1. Create a cloned voice using Instant Voice Clone
  2. Provide ~10 seconds of audio (extracted from the first generated video)
  3. Set language/accent labels (subtitles mention Spanish and “Mexican accent” due to similarity)
  4. Generate the “influencer voice”

CapCut role (for timing and extraction):

  • trims video segments to remove silent gaps (guided by audio waveform activity)
  • exports audio from the edited video

6) Audio replacement + preserve ambience (CapCut editing)

The creator replaces the original speech with the ElevenLabs cloned voice while preserving original ambient/background noise (e.g., wind/white noise).

This is done via CapCut audio separation features:

  • “separate audio”
  • “separate voices” (to remove original voice while keeping background sound)

The goal: make the result feel more real rather than overly clean.


7) Finishing polish: exposure/lighting, subtitles, styling (CapCut)

Final styling is handled in CapCut:

  • adjusts video look:
    • lowers exposure slightly
    • raises lighting a bit
  • adds subtitles:
    • generates subtitles in Latin American Spanish
    • selects subtitle templates and a font (mentions Poppins)
    • styles them (e.g., uppercase, font size changes)
    • re-aligns subtitle tracks on the timeline

Final Result Claims (Review)

The finished concatenated video is praised for:

  • perfected voice consistency across all three segments (described as identical voices in all clips)
  • being convincingly real, to the point that many viewers may think it’s a real recording

Tool stack explicitly summarized:

  • Google Flow + Google Omni / “UFO” for video
  • GPT / OpenAI image 2 for influencer image + product cleaning
  • ElevenLabs for voice consistency
  • CapCut for editing and subtitles
  • mentions Cloud was also used for generating one HTML-related prompt

The description also mentions free resources, including an “HTML prompt” download.


Main Speakers / Sources

  • Main speaker: the tutorial creator (speaks in first-person throughout, as the primary narrator/host)
  • AI systems referenced:
    • Google Flow (Google Omni / “UFO”)
    • GPT chat with OpenAI image 2
    • ElevenLabs (voice cloning)
    • CapCut (video/audio editing and subtitles)

Original video