Video summary

Cách làm VIDEO AI VEO 3.1 ĐỒNG NHẤT Voice & Nhân Vật | Tư duy tạo Prompt làm Video liền mạch 2026

Main summary

Key takeaways

Technology

Overview

This video is a Vietnamese, tutorial-style walkthrough on how to create AI videos using Google VO3 (VO3 Ultra). It focuses on a prompting workflow for:

  • Character consistency
  • Voice synchronization (consistent voice acting across clips)
  • Assembling short segments into a coherent narration video

It also includes a risk analysis related to YouTube monetization when publishing content that is fully AI-generated.


Key Platform / Product Features (VO3 / VO3 Ultra / Related)

Voice synchronization / consistent voice acting (new feature)

  • Described as a recently released capability that helps keep the same voice across multiple generated clips for the same character.

Ultra package benefits

  • Unlimited copies / bulk creation after paying monthly.
  • Ability to generate many videos within the same subscription window, with credit usage behavior controlled by a setting.
  • Requires selecting a lower priority / “Fed Lower” option to avoid using credits, so you can generate many videos without “losing money/credits.”

Credit / asset behavior

  • Video creation can be done without credits under certain settings.
  • Mentions that image credits may not be consumed depending on the chosen method/model (e.g., “banana 2” referenced as best).

Scene generation options

  • VO3 can generate videos from:
    • Start/end frames
    • Individual elements
  • Includes orientation choice (vertical/horizontal) and multi-image creation.

Character consistency via reference images

  • To keep the same character in talking clips, VO3 needs a reference portrait image.
  • Without a proper reference, the system may generate a different person/character.

Core Tutorial Workflow (End-to-End Process)

1) Split the script into short segments

  • The creator uses the idea that 8 seconds = one generated unit.
  • For a ~2-minute output, they suggest roughly 12 segments/commands.
  • Scaling idea: for longer videos, you effectively multiply 8 seconds by the target duration, generating a similar number of segment-units.

2) Write/generate “12 prompts” (12 commands)

  • Generate about 12 prompts/commands using an LLM (mentions GPT chat / “GBT” / “Gaminy”).
  • The prompts are designed so AI outputs visual scenes without text.

3) Avoid text inside the generated video

  • The speaker claims VO3 may:
    • crash
    • misspell Vietnamese text, especially with diacritics

4) Generate image scenes first (or generate short clips directly)

  • Create 12 descriptive image scenes first (or directly generate 12 short videos).
  • Character setup:
    • Upload the creator’s own photo (portrait-based identity).
    • Instruct VO3 to create a 3D/animated Hollywood-style character matching them.
  • Recommendations:
    • Use vertical aspect ratio for short-form content.
    • Use landscape for longer videos.
    • Maintain the same character identity across all scenes.

5) Convert each scene into a video clip

  • For each scene:
    • Copy the matching prompt/command
    • Provide the reference/portrait image
    • They emphasize checking selections (they compared two versions).

6) Enable voice synchronization for talking segments

  • Choose a voice and enable sync so the character speaks consistently across clips.
  • Limitations noted:
    • Harder to reliably express emotion changes (e.g., “excited vs sad”) unless prompts are detailed.
    • Voice can still vary in accent (Northern vs Southern), even when requesting Vietnamese accent.

7) Assemble the final video

  • Import the generated clips into a new project timeline.
  • Drag each clip into the correct segment region to match the script flow.
  • They recommend export sharpening/sharpness settings rather than brute force upscaling to 1080p/4K.

Practical Prompt Guidance (What They Emphasize)

Prompt purpose / narrative beats

  • The “philosophy” is: “Why are you executing these 12 commands?”
  • Prompts are meant to correspond to time-sliced narrative beats.

Scene-to-meaning matching

  • Using only one static image or generic prompts leads to boring videos.
  • Create multiple scenes that align with the script’s meaning over time.

“No text” instruction

  • Repeating the rule: include “no text” in prompts to reduce errors/crashes.

Character authenticity & branding

  • They advise using your own image to create a chibi/animated version for branding.
  • Avoid randomly generated AI faces that don’t match you.

Example Demo Theme: “Saving Money”

The tutorial demonstrates with a “saving money” theme:

  • Money drained by small daily expenses (water, taxi, outings, stress meals/drinks)
  • Need to distinguish necessary vs unnecessary spending:
    • Necessary: rent/electricity
    • Unnecessary: lifestyle treats
  • Ends with: “freedom comes from savings”

They also use a mini-game-like mapping method to match:

  • which prompt/scene number corresponds to
  • which part of the script.

Mini-game / Organization Method (To Reduce Mismatch)

  • Return to a Mini Games section and give a command to create a table mapping prompt scenes to video segments.
  • The output helps decide which quote fits which generated clip (e.g., scene 1 ↔ segment 1, etc.).
  • Presented as a time-saving consistency tool.

Risk Analysis: YouTube Monetization Concerns

Reported issue

  • They claim a fully-AI channel (example mentioned: Cường’s channel) was flagged for “dishonest content”.
  • Monetization:
    • Works in month 1
    • Then gets flagged again in month 2
  • Alleged trigger: repetitive AI video patterns

Advice / caution

  • If you scale too aggressively and violate platform rules, you may:
    • lose monetization
    • waste time and money
  • They suggest adding captions/structure and avoiding relying entirely on AI.

“Tool vs VO3 account” Explanation (Bulk Automation)

  • No external tool can fully replace VO3 generation quality.
  • The “tool” mainly helps connect to the VO3 account and speed up bulk operations.
  • Cost mentioned: ~149,000 VND/month
    • Not recommended for creators making short videos who already get good results directly in VO3.
    • More reasonable for YouTube creators who need bulk generation.

Account Types Mentioned

Genuine personal Ultra inside a Google Family

  • A family group (max ~5 people) using one Ultra subscription.

Reseller / provided accounts

  • Credits may be obtained quickly, but:
    • higher chance of being flagged by Google
    • higher chance of restrictions
  • Therefore, considered risky.

Additional Idea: LLM Translation / Summarization Workflow

The tutorial also mentions using an LLM to:

  • take subtitles/transcripts from foreign videos
  • translate/summarize them with GPT
  • rewrite Vietnamese content while preserving the original meaning

Main Speakers / Sources (Inferred from Subtitles)

  • Primary speaker:Cường” (repeated reference; likely the presenter and/or a creator used for comparison)
  • VO3 / Google ecosystem: VO3 Ultra features and the VO3 account used
  • LLM mentioned: GPT / GBT chat (used to generate commands/prompts and rewrite content)
  • Other referenced tool/service: “Gamini / Gamini” (mentioned as assisting with generation; exact identity unclear from subtitles)

Original video