← Back to video archive

Airdroplet AI summary

AI Video Just Got WAY TOO REAL... (VEO 3)

May 21, 2025Wes RothAI score 10028,082 views

Watch original on YouTube ↗

AI-generated summary

Here's a summary of the video about Google's Veo 3 AI model:

Google's new Veo 3 AI video model is seriously impressive, especially because it can generate not just video, but also integrated audio like music, sound effects, and even speech directly from text prompts. The video showcases a bunch of different, often wild, prompts to see how well Veo 3 performs, highlighting its strengths in motion, detail, and audio while also pointing out some inconsistencies and glitches.

Here are the key topics and technical details discussed:

  • Veo 3's Big Deal: Integrated Audio: The standout feature of Veo 3 is its ability to generate video with audio elements like music, voices, and sound effects based on the prompt. This is a major step up and makes the videos feel much more complete and dynamic compared to video models that only generate visuals.
  • Testing Methodology: All available AI credits were used to generate multiple versions (usually four) for a variety of prompts. The results shown are not cherry-picked; they represent the range of outputs from these tests, giving a realistic view of the model's performance.
  • Inflatable Duck Chase: Testing a "dirty off-road buggy racing through mud getting chased by a large, scary looking blow up duck." The results were phenomenal, especially the motion and menace of the duck. Version four was considered the best, even showing the duck gaining on and knocking the buggy off the road, which was surprisingly effective.
  • T-Rex Reflection: Trying to generate "two women slowly raise a mirror so that you can see your own reflection. You are a menacing T-Rex with massive teeth." The model handled reflections well, creating realistic-looking scenes, although there was some variation in quality across the different versions. Version one was personally felt to be the best overall.
  • The Hacking Octopus & Wet Keyboard: A longer, multi-part prompt about an octopus hacking a computer, hiding when someone enters, and the person asking "Why is my keyboard all wet?". The model did a great job with the multi-step narrative and capturing human expressions reacting to the wet keyboard. However, there were visual glitches like headless octopuses in some versions. A surprising and weird detail was that one generated scene looked uncannily like the presenter's actual keyboard setup.
  • Gorilla vs. 10 Men: Testing a "gorilla fighting 10 men" to see how the AI handles chaotic battle scenes. The results were pretty good, capturing the action and intensity. Version three was potentially the best, despite a slightly silly sound effect at the end.
  • First-Person Forest Run: A prompt for a "first person view of an animal running through a night forest with superhuman speed, eventually emerging to see a human village and people fleeing in terror." Most versions didn't quite capture the requested first-person view or the animal running correctly, but one version did it perfectly and was considered "really good" and by far the closest to the prompt's intent.
  • Eagle Playing Accordion: An absurd prompt asking for an "eagle... playing the accordion." The model generated different interpretations, some with human-like hands or extra limbs, which was weird. The audio felt appropriate, and one version specifically seemed to capture the "struggle" an eagle would have with the instrument.
  • Undead Guitar Solo: Asking for an "undead from Dungeons and Dragons is playing a guitar solo on top of a mountain of skulls. A field of skeleton fans are going wild down below. The moon is bright and red." This prompt highlighted the AI's ability to generate music on the fly to fit the description. The visuals were detailed, capturing the undead look and the scene effectively, even adding some ad-libbing sounds in one version.
  • Yarn Sumo Trash Talk: Generating "two sumos made out of yarn... doing a playful trash talking" with specific lines provided. Despite a typo in the prompt ("Yarm"), the AI understood "yarn" and generated characters that looked like they were made of yarn. The speech fidelity and lifelike gesturing in some versions were impressive, though one version had disturbing visuals and another wasn't clear who was speaking. Version one was considered the best overall for this prompt.
  • Wolf Chasing Rabbit: A "first person view of a wolf chasing down a rabbit, jumping over falling trees and branches... View low to the ground." Similar to the forest run, some versions weren't strictly first-person but still captured the speed and feeling of the chase effectively. Version three was particularly liked for capturing the intensity.
  • Walking Brick House: Prompting for a "brick house with people leaning out of windows. It has six mechanical legs and is walking down the street as people stare in awe." Version one was the most realistic, showing people leaning out and looking like a walking building. Other versions looked a bit "off," highlighting the inconsistency in rendering complex, unnatural concepts.
  • Fat Cat on Throne: Asking for an "obnoxiously fat cat sits upon a large golden throne. It looks at you as you approach and says, I see you brought me snacks. I guess I will let you live for meow." The model successfully generated the scene and synthesized the requested speech in three out of four versions, even adding a cat pun. One version captured the attitude but failed to deliver the specific lines. Version one was felt to be the best.
  • Spaceship Approaching Ring World: Attempting a notoriously difficult prompt: "a view from the cabin of a spaceship as it approaches a massive ring world... Signs of a civilization can be seen on the inner part of the ring world." As expected, the model struggled with the 'ring world' concept, which AI models typically find hard. While none were perfect ring worlds (some looked like Saturn's rings), version three was the closest and considered among the best renditions seen of this specific challenging prompt.
  • Revisiting Veo 2 Prompts: Testing prompts previously showcased by Google for Veo 2, like the "first person chasing ice skater" and the "helmet mounted POV tailing a woman on a dirt bike." Veo 3 handled these well, with excellent motion and capturing the requested points of view, also adding great sound effects like the ice skates or dirt bike noises.
  • Roller Coaster POV: A "first person view of a slowly rising roller coaster before it drops rapidly into the night below." The model captured the scene beautifully, including the stars, but consistently failed to include the requested "drop" portion, cutting off right before the climax.
  • Snow Tiger: Generating a "tiger made out of snow, walking in a snowy forest." Some versions captured the 'made of snow' look perfectly and had fantastic sound effects like crunching snow (rated A+), while others looked more like regular tigers in snow or had less fitting sounds. The variation in results was apparent here.
  • Overall Impression: The presenter was very impressed with Veo 3, particularly the quality and integration of the audio (sounds, music, speech, intonations). He felt he ran out of credits too quickly just as he was learning how to prompt it effectively, suggesting that mastering prompting is still key. He believes the model is "very, very good" and possibly represents the "next generation" of AI video models.
  • Actionable Takeaway: There's a clear need to get more credits and continue testing to better understand how to prompt Veo 3 to get the best results, indicating that effective prompting is a skill that needs development even with advanced models.

Video transcript

Open transcript
Why is my keyboard all wet? So the new VO3 model is out and it's been absolutely blowing my mind. It's really good. It has music added. It has voices added. It has sound effects added. Whatever audio you want to add to your video, it does it. It does in the prompt. You just type in what you want it to say and it just goes for it. So here I went through all of my AI credits that I had with VO and generated a bunch of different prompts that I wanted to see how well it would perform. So here are basically all of them. I think maybe there's like one or two I left out, but these are not cherry picked. These are not like the best of the best. These are just all of them. Let's take a look. This is a dirty off-road buggy as racing through mud, getting chased by a large, scary looking blow up duck. All right, let's see. This is version one. That's pretty menacing. That's a pretty menacing duck how twaddling after that truck. I just absolutely, absolutely phenomenal. All right, here's version two. Wow, we got some air there. The motion of the duck is just phenomenal. You can tell it's a large inflatable thing. Wow. Here's version three. Still very good. Oh wow. It kind of bypasses it. Phenomenal. And we have version four. It's gaining on the truck. It's about... Wow, it knocks it off the road. I think that was the best one yet. Everything about that is just looking absolutely phenomenal. And kind of scary too, I gotta say. They really captured the prompt there. All right, next I want to see how well it does reflections. So, two women slowly raise a mirror so that you can see your own reflection. You are a menacing T-Rex with massive teeth. All right, so here's a version one. So, it's looking very real. Great reflection. Pretty good. Here's version two. I feel like version one was better. The reflection was better, but otherwise it's excellent. Here's version three. They're lifting the mirror and... That's pretty good. And version four. They're all great. I gotta say, I feel like version one was the best one. Just everything is perfect. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Oh wow. Alright, this one didn't come out perfect, but still there's a lot of great things here that I think we all should see. An octopus climbs out of its tank to try to hack a computer. When he hears someone coming, he rapidly climbs back into his tank. A person walks in and asks, why is my keyboard all wet? So notice that's kind of a long-ish prompt, but let's take a look at what we generated. So this is one. Why is my keyboard all wet? That's terrific. Her expression. Why is my keyboard all wet? It's phenomenal. Here's version two. It just pops on the keyboard. Wow. Why is my keyboard all wet? The captions need work. But other than that, there's a lot of great things here. Why is my keyboard all wet? These are excellent. Unfortunately, here he's not really in the tank. He's like, oh, I'm not going to do this. I'm not going to do this. I'm not going to do this. I'm not going to do this. I'm not going to do this. I'm not going to do this. Why is my keyboard all wet? Why is my keyboard all wet? Why is my keyboard all wet? This is three. Why is my keyboard all wet? Why is my keyboard all wet? Why is my keyboard all wet? Why is my keyboard all wet? This is another One. Why is my keyboard all wet? Why is my keyboard all wet? Halfway out of the tank, but still. Here's another one. This is another Headless Octopus? Why is my keyboard all wet? I think this is the greatest react. Human reaction, at least. Why is my keyboard all wet? I love that face. Like, how did that happen? And the octopus that jumps is, you know, not great because we're missing the head, it seems like. But in this initial shot, this octopus, I mean, it's perfect. So I don't think any of them are perfect fidelity, but I got to say there's so much good stuff happening here that I just got to give it, I got to give it points. What's weird is like, this is literally my keyboard. I'm pretty sure this is a Razer mouse, the same one that I have. And this looks a lot like my computer. This looks nothing like my monitor, though. All right, how about gorilla fighting 10 men? Let's see how well it sort of is able to do a chaotic battle scene. Let's try the first one. It's pretty good. Number two. Wow. Scary. Here's a three. Ouch. This one might be the best one yet. That little sound effect towards the end is kind of silly, but other than that, it's very, very good. Here's a first person view of an animal running through a night forest with superhuman speed, eventually emerging to see a human village and people fleeing in terror at the sight of it. So in this one, I already know that only one of them did well. So let me show you the ones that didn't do well. Okay, so that's okay. That's all right. None of these really captured what I'm asking for as an animal running through the forest. Except I think number one does it perfectly. That one was by far the closest and really good. Have you ever wondered what an eagle would look like if it was playing the accordion? You know how it's got those razor sharp claws? You've wondered about that too, I'm sure. Well, here it is. Version one. I mean, the sound sounds good, right? Let's see. Number two. This I feel really captures the struggle that the eagle would have pushing the buttons, you know, accurately. Here's three. This one's the best accordion player, but these are human looking hands. And this one has an extra hand here. Not sure what's happening there. All right, how about an undead from Dungeons and Dragons is playing a guitar solo on top of a mountain of skulls. A field of skeleton fans are going wild down below. The moon is bright and red. All right, let's check it out. I mean, I like a lot about that, especially this little up close shot where you really get to see the undeadishness, I guess. Very, very good. Two. What blows me away is like it's generating the music on the fly, just based to just to fit the description. Here's three. Yeah. Oh, yeah. You guys are awesome. Yeah. That's kind of towards the end there. Kind of added some little ad-libbing, but still very, very good. Here's four. Yeah. I don't know. I think it's between one and four for me. So here I wanted to have two sumos made out of yarn. As I'm saying this, I realized I spelled it Yarm. I don't know what Yarm is, but I think it understood I meant yarn. So they're preparing to fight and they deliver. They're kind of like doing a playful trash talking. And I wrote out what I wrote out what they should say. Take a listen. My highlight reel has you in every frame face down. Your belt is the only thing in this ring that still thinks you can hold something. I really liked the second one, sort of his little like gesturing here. It feels very, very lifelike. My highlight reel has you in every frame face down. Your belt is the only thing in this ring that still thinks you can hold something. It's pretty good. No background, but I like the little characters. Here's three. My highlight reel has you in every frame face down. Your belt is the only thing in this ring that still thinks you can hold something. That's not too good. Just because you can't really tell if the second person is talking that the one on the right is talking. My highlight reel has you in every frame face down. Your belt is the only thing in this ring that still thinks you can hold something. This one, I think, has the highest sort of voice fidelity, but I'm going to say the most disturbing visuals. I got to give it to one. I think one was the best. My highlight reel has you in every frame face down. Your belt is the only thing in this ring that still thinks you can hold something. Yeah, I think that one nailed it. Here's a first person view of a wolf chasing down a rabbit, jumping over falling trees and branches. The rabbit darts left and right, trying to escape. View low to the ground, so you feel the immense speed of the chase. Here's one. Here's two. I really like that one. Here's three. I like it a lot. These two aren't first person, but they really captured the feeling that I wanted. Here's four. So not quite what we're looking for, but I got to say, I mean, this one really, I think really good. Really captures that chase. Here's a brick house with people leaning out of windows. It has six mechanical legs and is walking down the street as people stare in awe. Here's one. Here's two. Here's three. Here's three. And four. All right. So, I mean, I guess one is the best. You can actually see the people up there kind of like rocking back and forth. I mean, this looks real. The rest of them looks a little bit off, I would say. So interestingly, this actually didn't render the first time I ran them, but now I have all four. In the beginning, I only had this one. So actually, I'm seeing these for the first time. The prompt is an obnoxiously fat cat sits upon a large golden throne. It looks at you as you approach and says, I see you brought me snacks. You know what? I'll, I'll, I'll have the cat say it. Here's one. I see you brought me snacks. I guess I will let you live for meow. That's pretty good. Two. I see you brought me snacks. I guess I will let you live for meow. Terrific. Three. Meow. Meow. Meow. Meow. It didn't translate from cat. Meow. Meow. I mean, capture the attitude, but didn't deliver the line. Four. Four. I see you brought me snacks. I guess I will let you live for meow. Terrific. I feel like one here takes the, uh, takes the cake. And here's one of the hardest prompts that no model is able to do. I haven't seen any good renditions of it. So it's a view from the cabin of a spaceship as it approaches a massive ring world. Anything to do with a ring world is just never well rendered. A giant structure in the shape of a ring that rotates around the sun. This word, if you're wondering, is supposed to be signs. Signs of a civilization can be seen on the inner part of the ring world. One thing I like about AI is it always gets what I'm trying to say. Here's one. I mean, that's not quite ring world, but, um, it's good. I mean, you can tell it's a massive structure. You can see the little details on the surface. It's, it's good. Here's two. Yeah. I mean, there's definitely something magical about it. So these are like the rings of Saturn. So still not quite what we're looking for, but I mean, I'm liking what it's doing. Here's three. That one is the closest. Again, keep in mind, I haven't seen anything rendered this perfectly, but these are some of the best ones I've seen for sure. And it didn't do a four for some reason. I think this one's my favorite. Here's a continuous first person shot captures us chasing a woman ice skating across a vast, glassy frozen lake surrounded by snowy peaks. This is one of the ones that Google was showcasing from VO2. So I just wanted to see how VO3 would handle it. So that's pretty good. That one's excellent. You can definitely hear that the ice skates on ice. It's great. Good. And here's four. Great sounds. I gotta say, like really kind of captures that. So here's another one that VO2 did. So I just wanted to see how well it does it here. A continuous helmet mounted POV shot shows us tailing a woman on a dirt bike as she races across rolling desert dunes. Here's one. That one's a bit weird. That's a little bit weird. That's a little bit weird. Okay. Here's two. That's pretty good. You can see them kind of getting some air. Very cool. This is three. Yeah, a lot of great stuff happening here. And four. Yeah, these are all very, very cool. Definitely nails the prompt. This next one is a first person view of a slowly rising roller coaster before it drops rapidly into the night below. Here's one. That is pretty good, I gotta say. Here's two. Very cool. This is three. I love how it captured the stars, like this whole thing. It looks phenomenal. I love it. But there's, this is not a drop. Like if it dropped right here, it would have been perfect. This is more like a flat sort of straight away. Yeah. I wish there was a drop that this, that would have made this perfect. Here's four. That one's really good. But again, like it cuts off right before the drop. This would have been great with the drop. And here's tiger made out of snow, walking in a snowy forest. Here's two. Really good. I love the, I love the sound of the snow. It's so perfect. Here's two. So here it went with no sounds, more of a, like a single note, but like the tigers look phenomenal. They look like they're made out of snow. These two, I mean, this one, you can tell it's made out of snow. This one might looks like a snow covered tiger. No, all right. That's okay. So here's the, I'm looking for here's four. All right. So something weird was happening in these two, but man, this one is phenomenal. Listen, let's listen to this. The crunching of the snow. This one's like an A plus for, for me. All right. So I blew through all of my credits. Uh, I might re up tomorrow and do some more testing. I'm very impressed with it, especially the sounds and the sounds and the music, the speech, the intonations. So much of it is so good. And I feel like I ran out of credits too quickly, just as I was getting to start to understand how to really prompt it properly. So I'm definitely going to get some more credits and, um, and run this again in the future because this model is very, very good. But let me know what you think. How was the sound? How was the music? How was the, the, the graphics, how, how it renders different scenes. Does this feel like kind of the next generation AI video models, or are you still unimpressed? Well, let me know if you made it this far. Thank you so much for watching. My name is Wes Roth and I'll see you next time.