How I Built an AI Avatar and Hype Video in 15 Minutes with Google Flow
I put Google Flow and the new Gemini Omni model to the test by creating an AI avatar of myself and producing a full podcast hype video from scratch. Here’s the entire workflow, from scanning my face to the final, slightly uncanny result.
Claire Vo
Full episode
Watch or listen
Workflows from this episode
- How to Create a Promotional Video with an AI Creative Director
- How to Create a Personalized AI Avatar with Google Flow
Episode outline
I built an AI version of myself, generated seven video scenes, and cut together a one-minute promo for How I AI in roughly 15 minutes. The surprising part was not the speed. It was how much of the creative process Google Flow handled for me, from brainstorming shots to assembling a usable first cut.
In this episode of How I AI, I tested Google Flow and Gemini Omni as someone without a video-production background. The question was practical: could these tools handle enough of the directing, prompting, editing, and scene planning that I could make something genuinely shareable instead of another AI demo that falls apart outside a screenshot?
The result landed somewhere between impressive and deeply strange. Some clips looked uncannily like me, especially side profiles and calmer shots. Others drifted into obvious AI territory with changing hair, fake glasses, awkward expressions, and generic sci-fi overlays. The workflow itself worked. Character consistency, emotional realism, typography, and visual continuity still required a human eye.
Building a reusable AI avatar from a phone scan
Flow’s avatar feature creates a reusable character from a guided phone capture. I had already tried the feature once when it first launched and failed to get a usable result, so part of this experiment was simply checking whether the setup experience had become more reliable.
Capturing the avatar
- 1. Start: In Flow, choose Create an avatar and launch the guided setup flow. See How to Create a Personalized AI Avatar with Google Flow.

- 2. Connect the phone: Scan the QR code on desktop to open the capture tool on your phone. Flow requests camera access and walks through the process from there.
- 3. Capture the face: The system asks you to look at changing numbers on screen while holding still, then rotate your head left and right so it can capture multiple angles and facial geometry.
- 4. Wait for processing: After the scan uploads, Flow processes the capture into a reusable character that can be referenced inside prompts and storyboard scenes.
- 5. Review the result: The finished avatar was recognizable immediately, though it had a slightly distorted wide-angle look. Some facial details felt accurate while others shifted depending on the camera angle.

Using Flow as an AI creative director
The interesting part of Flow is not just the video generation. It combines planning, prompting, character references, scene generation, and editing into one interface. In practice, it behaves more like a lightweight AI production suite than a standalone video model.
Developing the storyboard
I started by asking Flow’s chat agent to help create a hype video for How I AI using the saved avatar version of me as the main character.
The initial prompt was intentionally simple:
Example storyboard prompt: Help me create a short hype-video storyboard for the How I AI podcast using the avatar I already made.
Example context: The show is hosted by Claire and focuses on practical ways to use AI at work and in life. Propose a small set of scenes that will make the premise immediately clear.
Instead of jumping straight into generation, the agent asked follow-up questions about the environment, pacing, mood, and overall visual direction. That interaction mattered because it shifted the process from prompt vending machine to collaborative planning.

I described a dark green home office with AI books, posters, practical lighting, and what I called a more authentic lifestyle version of a hacker aesthetic. I wanted something grounded enough to feel like a real creator workspace, but still cinematic and technical.
Example scene direction: Put the avatar in a dark home office with deep green walls, AI books, playful posters, and practical lighting. The room should feel lived-in and personal, with a high-tech coding and hacker edge.
Flow responded with a seven-scene storyboard. The sequence included extreme closeups of typing on a mechanical keyboard, a wider office reveal, a dramatic chair spin toward camera, a floating heads-up-display effect, and a podcast call-to-action ending. Some concepts leaned heavily into stock AI imagery, but structurally it was useful. The system produced a coherent shot list with pacing and progression instead of isolated prompts.
What stood out most was how quickly the brainstorming loop collapsed. Normally, planning a promo like this would require references, rough scripting, shot ideation, and some understanding of editing rhythm. Here, the tool handled a surprising amount of that creative scaffolding.

Generating scenes with the avatar
Once the storyboard looked usable, I copied individual scene descriptions into the generation interface and replaced references to my name with the @me avatar tag so the model would use the saved character.
My first generation failed for a very human reason: I accidentally left the mode toggle set to image generation instead of video generation. The result looked fine as stills, but obviously defeated the entire point of the exercise. Switching the mode fixed it immediately.

For the chair-turn reveal scene, I used the storyboard language almost directly as the prompt:
Example performance direction: I spin my ergonomic chair toward the camera, push my glasses up the bridge of my nose, and say, "I am Claire, and this is How I AI."
Each prompt generated two video variations. That mattered because quality fluctuated dramatically between takes. One version captured my hair and face more accurately. Another added fake glasses and exaggerated movement. One clip unexpectedly inserted robotic narration that I never asked for, which became a good reminder that audio outputs need as much review as visuals.
The generated scenes also inherited details from my original environment in unexpected ways. Posters and books from the room behind me during the avatar scan reappeared inside generated scenes, even when the office layout changed. At times that helped realism. At other times it created continuity problems because every clip interpreted the same room differently.
The strongest shots were surprisingly subtle ones. Side profiles while typing looked convincing. Slow turns toward camera held together reasonably well. More expressive moments broke much faster. A laughing shot landed directly in uncanny-valley territory with unnatural facial timing and exaggerated expressions.
Flow also leaned heavily into generic cinematic AI symbolism. The model repeatedly inserted glowing interfaces, oversized futuristic displays, floating schematics, and dramatic HUD overlays. In one shot I appeared to be inspecting what looked like a giant digital church blueprint on a massive transparent tablet. None of that was requested directly, but it revealed the model’s visual priors around "AI creator" aesthetics.

Editing the promo inside Flow
Flow includes a browser-based timeline editor, which meant I could finish the experiment without exporting everything into separate editing software. That reduced friction more than I expected.
I selected my preferred take from each generated scene, arranged them according to the storyboard order, trimmed the clips, and adjusted transitions. The assembly process took about five minutes. Most of the work was not technical editing. It was deciding which generation felt the least uncanny and which continuity mistakes were easiest to ignore.
That selection process became the real creative role. The AI could produce many candidate clips quickly, but someone still needed to judge pacing, realism, tone, and whether the avatar actually resembled the intended person from shot to shot.

Reviewing the finished video
From the initial avatar scan through the final export, the complete promo was assembled in about 15 minutes. That included figuring out the interface, generating multiple scenes, regenerating failed outputs, and editing the final cut together.

I considered the final result shareable, but not polished professional work. The biggest breakthrough was not visual perfection. It was the ability to move from zero concept to a coherent promotional sequence without specialized production skills.
The finished video genuinely resembled a rough creative draft a small team might produce early in a campaign process. It had structure, mood, recognizable visual motifs, and a usable narrative arc. That alone changes who can experiment with video creation. You can follow the full implementation in How to Create a Promotional Video with an AI Creative Director. See How to Create a Promotional Video with an AI Creative Director.
What worked surprisingly well
- Speed and accessibility: The integrated workflow collapsed brainstorming, generation, and editing into one loop. I could test ideas immediately instead of spending hours planning before seeing results.
- Character resemblance in calmer shots: About half the generated footage looked convincingly like me. Side-profile shots and less expressive moments were consistently stronger than direct emotional performances.

Where the illusion broke down
- Continuity problems: Hair length, wall colors, books, plants, lighting, and office layouts changed from clip to clip even though every scene was supposed to take place in the same environment.
- Emotional realism: The avatar struggled most with expressive moments. Laughing, exaggerated movement, and conversational delivery exposed the uncanny quality quickly.

- Default AI aesthetics: The system repeatedly introduced generic futuristic imagery like glowing interfaces, dramatic HUD graphics, and oversized holographic displays that felt disconnected from the grounded creator-office setup I described.
- Typography and graphics: The generated text overlays and end cards were noticeably weaker than the video footage itself. Motion and cinematography advanced much faster than branded graphic design inside the workflow.
How I would approach a second version
On a second attempt, I would reinforce the environment description in every single prompt instead of assuming the model would remember it across scenes. I would also add stronger visual references, define typography explicitly, and budget for multiple generations per shot rather than expecting the first result to work.
I would probably narrow the emotional range too. The tool currently performs better with composed, cinematic actions than expressive acting. Leaning into that limitation would likely produce a more believable final sequence.
The biggest gain here was creative leverage. Flow acted like an AI producer that could brainstorm shots, draft sequences, generate options, and hand me editable material fast enough that experimentation became cheap. That changes the psychology of making video. Instead of investing heavily before seeing anything, I could iterate in real time and react to actual footage.
This workflow is worth copying if you need rapid concept videos, social promos, visual prototypes, or rough creative direction without a production team. The speed is real, and the integrated workflow lowers the barrier dramatically for non-editors.
What still requires human judgment is everything related to taste and trust: choosing believable takes, maintaining continuity, deciding when the avatar crosses into uncanny territory, and cleaning up the generic AI aesthetics these models still default toward. For polished brand work or identity-sensitive content, the AI can accelerate production, but it cannot yet replace a careful editor or creative director.
Watch or listen
Sponsors
Thanks for supporting How I AI
Prioritize with insights, build with confidence
Connective infrastructure for production AI
Build your next product with ChatPRD
Turn an idea into a PRD, user stories, and a plan.


