← Back to video archive

Airdroplet AI summary

AI images just got dangerously good (RIP diffusion??)

March 26, 2025Theo - t3․ggAI score 100221,572 views

Watch original on YouTube ↗

AI-generated summary

OpenAI just dropped a major update to its image generation capabilities within ChatGPT, powered by the new GPT-4o model. This feels like a significant leap forward, especially compared to their previous efforts with DALL-E, and potentially challenges other image generation tools like Midjourney, particularly in areas like text rendering and object accuracy. The most surprising thing is that it seems to be integrated directly into the core 4o language model, unlike previous diffusion-based methods.

Here's a breakdown of what's new and how it works:

  • A huge upgrade from DALL-E: OpenAI's previous image model, DALL-E, wasn't great. The presenter showed an example image of a programmer with weird text and blurry elements, noting that smaller companies like Midjourney had been far ahead in image quality. The new 4o image generation seems vastly improved in realism and detail.
  • Impressive initial performance: Trying the new model with the same prompt (JavaScript programmer), the output is much better. Skin textures are realistic, the mustache looks good, and the overall image is solid.
  • Better Text Handling: This is a major point. Previous image AI often struggled with generating legible or correctly placed text. The new model can generate text correctly on things like t-shirts (even getting the JS logo right) and diagrams. While it didn't perfectly code the Fibonacci sequence on a laptop screen, the attempt was much better than expected. It also handled changing text in existing meme images, even intelligently preserving parts of the original image underneath the new text, which is fascinating.
  • Reflection and Object Accuracy: A really surprising capability is the model's ability to handle reflections accurately, like a photographer's reflection in a mirror. It can also place images of people in front of surfaces like whiteboards and add text to them with decent results. The model is also much better at handling multiple objects and their relationships within an image, claiming to handle 10-20 distinct concepts compared to the 5-8 other systems struggle with. An example showed it correctly placing 16 specific objects in a grid based on a list.
  • Different Technical Approach: This is a big reveal. Unlike traditional diffusion models (like DALL-E or Midjourney) which start with noise and refine it over multiple passes (like refining a blurry image), GPT-4o's image generation is described as an "autoregressive model natively embedded within ChatGPT." It uses the same core architecture as the language model, allowing it to leverage its knowledge base and context in generating images. This is potentially why it's much better at things like text and following precise instructions. The presenter found this technical shift fascinating and wished more details were available.
  • Speed and Process: The generation process wasn't lightning fast during the demo, taking a noticeable amount of time. It also generates top-to-bottom in a way that looks different from typical diffusion models, which the presenter found weird given the diffusion explanation.
  • Character Consistency: The model seems pretty good at maintaining character consistency across multiple prompts, though it wasn't perfect (like changing a character's hair color in one test). This is useful for creating visual storyboards or consistent elements for projects.
  • Safety and Limitations: OpenAI is incorporating safety measures, including C2PA metadata to identify images from GPT-4o. They also have an internal tool to check if an image originated from their model (though this isn't public, which makes sense). They have heightened restrictions on generating images of real people, particularly for potentially harmful content, but seem more flexible than previous models about simply including real figures in non-offensive contexts (like Donald Trump eating ice cream).
  • Useful Features: Key practical additions include the ability to specify aspect ratios, use hex codes for colors, and generate images with transparent backgrounds. The transparent background feature worked really well in the demo, producing clean edges, which is something many other tools lack. Upscaling didn't seem to work correctly during the test, generating the same low resolution as the original.
  • Integration with Sora: The new 4o image generation capabilities are also integrated into Sora, OpenAI's video model. This means you can use 4o to generate a better starting image to then turn into a video. Sora can also generate multiple images at once, a feature missing from ChatGPT's one-at-a-time generation.
  • Implications and Takeaways: The improved text generation is a major win, making these tools much more practical for tasks where text is needed (diagrams, memes, etc.), potentially reducing the need for external editing software. The ability to handle complex prompts and multiple objects opens up new possibilities for using AI images beyond just decorative art. However, the presenter also expressed concern about the potential for spreading misinformation given the increased realism and ease of use.

Overall, this feels like a significant step for OpenAI in the image generation space, catching up to and potentially surpassing competitors like Midjourney in specific capabilities, driven by a potentially different technical approach. The precision, text handling, and integrated nature within 4o make it a powerful tool with exciting, and slightly scary, implications.

Video transcript

Open transcript
OpenAI is a company known for breaking new ground in the AI space, at least in the LLM space. Their ability to generate models that generate text is pretty much unmatched to this day. It's cool seeing more companies start to catch up, but OpenAI has always been really far ahead because they invented the large language model as we know it today. That said, there are other types of generation that AI can do that aren't necessarily the strongest suit for OpenAI, in particular, Diffusion. For like three or four years now, ImageGen has been kind of a weak spot for OpenAI, and smaller companies like Midjourney that are fully self-funded and just doing a Discord bot have run circles around OpenAI stuff. That is, until today, because not only has OpenAI leapfrogged and made image generation significantly better than it used to be, they also appear to have figured out text. This is super cool. I'm really excited to dive into all of the things OpenAI just shipped and the benefits we can see, not just in ChatGPT and 4.0, but also apparently in Sora. So this is going to be very, very interesting. That said, as a $200 a month pro user of OpenAI, this bill needs to be paid. So a quick word from today's sponsor, then we're going to dive right in. Today's sponsor is one of those products that I can't imagine running my business without. And I mean that. I've been using them since way before I was sponsored by them, because I reached out to them and convinced them to sponsor me, because I just really like the product. Of course, I'm talking about PostHog, the best analytics product ever made, and have tried all of them. I changed providers every three months for like most of my career in the entire history of Ping until I finally moved to PostHog, and then it just was good. I love their Why PostHog page. It really shows you the tone and the way they think of the company, but screw all of that. I want to show you what it's like to actually use. Here's my real dashboard for T3Chat, because T3Chat, like all of my other products, has all of its analytics run through PostHog, and it makes everything so much easier. We wanted to start getting better observability on our LLM stuff. Turns out they have a package where you can just wrap all your LLM calls and get much better information all about them. I can click this button, and you can see why we have to do ads, because this is our costs for the last seven days, and it's missing some data because I have to fix some stuff I broke. It's so useful having something like this that can help me track all of the chaos that is running a company with actual stuff going on in it. I've also been talking to more founders recently and investing in a bunch of companies from the recent Y Combinator batches, and the ones that are using PostHog are significantly easier to work with because they know much more about their users. If you aren't using PostHog, you're using one of the competitors and you're paying too much, and if you're not using one of the competitors, then you're missing a bunch of information you need about your users. PostHog is borderline essential if you want to actually understand what people are doing with your stuff. And more importantly than all of this, it's open source. Yeah, really, you can go host it yourself if you want to. I am blown away with these guys. I am so happy I started using them. You should try them out today at soydev.link slash posthog. So 4.0 image generation has arrived. Quick history lesson. Previously, if you wanted to do image gen in chat GPT, you would use Dolly, which was their image generation model. And the quality of the things you can get from it were not great to say the least. I asked it to generate me a JavaScript programmer with blonde hair and a really nice dark mustache. And here's where we got it's not like the worst, but the text on the wrong side of the screen, the weirdness in the background here, the blurry kind of like, I don't know how to call it. It's almost like, why is it a worse version of everyone's favorite font? You know, I don't want Comic Sans in my AI gen. It's just the more you look, the worse it gets. And remember, guys, Midjourney hasn't raised a whole bunch of money and spent billions of dollars on training. They're a small team that has managed to make models that are significantly better than what OpenAI has shipped, at least until now. So what's the performance like with this new model? Let's try it. 4.0. Hit the Create Image button. JavaScript programmer with blonde hair and a nice dark mustache. Let's see how it does. First thing you'll notice is it is not particularly fast. Here we see still going, it's going to take a bit. We'll do our best to not edit so you get the real time speed here. Start at 1.32.15 or so. Another interesting thing you'll see here is it's generating top to bottom. You understand how diffusion models work. This is probably confusing to you as it is to me. I'll do a quick demo of how they work in a moment. But we can already see the results here are significantly better. It just changed the color halfway or so through the generation. I've noticed doing that consistently. Oh, no, don't put the code on the wrong side of the laptop again. Oh, it's the laptop is backwards. Better. Improved. Significantly improved. Mustache looks pretty solid. Nice curl going on it. The skin textures are insane. To get skin that realistic with an AI model is nuts. This is solid overall. The text on the laptop isn't great. Change the laptop's text to show a demo of the Fibonacci sequence in JS. Apparently, it can do text changes really well. Just to show the first one that we did before, it did a great job. It even has the right JS logo on the shirt with the font and everything proper. It does text way better than things I've tried in the past. It's possible that the refresh there just killed this in progress. I would give them crap for it. But our UI handles refreshes mid generation really badly to something that we're all doing wrong because resumable streams is a very hard problem to solve. I have one of my friends at OpenAI here. It should be better quality because it's the same model as the rest of the text. Yep. It really is just using 4.0. That's super cool and fascinating. I am curious if it's using 4.0. Is this diffusion still? How did they make this? Because this feels like it couldn't be diffusion in the traditional sense if it's 4.0. While this generates, I want to give a quick summary of the difference between these things. The way that LLMs work, to show you the basic original example, it's easiest to think of an LLM almost like really good autocomplete. Like if I type sup blank, if you know me well enough, you know the thing that's most likely to follow here is nerd. If I regularly ask my partner, what do you want for dinner based on previous chat histories, previous things I have done on my phone, it's smart enough to use that context to guess what word is most likely. It might have different options. Like it might have breakfast, lunch, dinner, laundry. I don't know. But it has these different options. And it has to decide based on the context we have here, what is most likely to be next between these different options. And that's why when you generate text using traditional LLMs, it comes in token by token. It is effectively breaking this up into chunks it can understand and using crazy math to figure out what chunk is most likely to be next. This is how standard LLMs and generative text work. This is not how diffusion works at all. So get this out of your mind. Diffusion is an entirely different thing. I'm going to show how it works a weird way. There was actually recently a new text model called Mercury Coder, which is a diffusion language model. And you'll see on the left here, we have a traditional LLM where it's generating token by token each piece as it goes. The way that diffusion works is it has a bunch of noise and it tries to correct the noise. If you're familiar with traditional image processing like sharpening, so we have this image that's four pixels and we tell an AI upscale it, what it will do is it will take this, it will cut it into the additional pieces. And then based on context clues and other higher res images, and most importantly, images that it has that are lower resolution that it is using as referenced by just downscaling and upscaling to generate a ton of data, it could do the rudimentary pass here where you just turn each pixel into its counterpart. So each pixel becomes the four of the grid blue, blue, blue, blue, blue, blue, blue, blue, blue. And then red here gets these four corner spots. This is the very rudimentary scaled upscale because we went from one pixel to four in each block and we just put the four in each of the spots. But let's say we knew this was part of something. We knew this was an American flag or we knew this was someone's shirt. We might be smart enough to do something like this, where now it's clear this is a line. And the model was smart enough to recognize that and basically make an upscale that isn't one-to-one. Because here we have the line drawn, but we don't have enough pixels to represent it in as fine-grained a way. But as we upscale, if we know enough about the image and what it's supposed to be, we can make these types of decisions as we upscale what we have. You think of a diffusion model as this ramped up a ton, where we're trying to take this thing we know is wrong and do a pass to remove some noise, do a pass to increase the resolution. Diffusion models are going through something we know is slightly wrong and trying to correct it as we go. And it does multiple passes because it's similar to like a noise filter. If you're trying to remove noise from an image, which is something we've been trying to solve for decades now in image processing, you can run the denoiser multiple times and it will remove more noise each time. Diffusion models are basically doing this. They are making corrections to a thing by passing over it over and over again with an instruction to try and make it look more like it. So when you use a tool like midjourney or any of these image generation tools, you say, generate me an image of a dog, it will start with a bunch of noise, a bunch of random gray dots, and then each pass will adjust them to look more like a dog. And if you use a tool like midjourney, image of a golden retriever, you'll see once it starts, you get an image. You couldn't even see how noisy it was because it's so fast now. But if I turn off the fast compute here, we'll do relax this time. I'll run it again. There. It gave us the really noisy early one, a phase if you can go back, pause, and zoom in on one of those so people can see. It starts with much noisier stuff. So if you start with this and you tell the model it's a picture of a cat, it can start to figure it out. And each step, each diffusion pass gets closer to what you want. Hopefully this all makes sense. This is also the reason that this text model that these guys Inception made is so bad because turning a bunch of noise and random text into functioning code is something we've all had to do if we've had bad enough interns, but it's not a pleasant experience. So for different tasks, different types of generation make more sense. And historically, OpenAI is much stronger on the left to right generation, token by token. And other companies have been able to go ahead of them with diffusion. You can also see why diffusion we've had a text in general and why we have the hilarious outputs that we got with the generation here with Dolly. This text being so awful makes sense if you remember that this started as noise and the noise has been refined to this point. Speaking of text, let's go back to here where I told to change the text of the code here. Did it change the laptop's text to show a demo of the Fibonacci sequence in JS? Let's see how it did. I didn't highlight it properly, but syntax highlighting is pretty hard. It's sizing things wrong. I don't know why there's a colon here. It tried its hardest. Didn't quite get it right, but it did try. Let's give it something more fun. Let's give it a meme. Give it that image. Change the text to be a meme about OpenAI image gen being better than mid-journey. And to be fair, we can give it to Gemini as well. I'm not sure if it'll even work in here. I'll throw it in Gemini as well. Why not? We really need to add image stuff to T3 chat. I've wanted to for a while. It can parse images, but it can't create them. It added new text there. OpenAI image. It didn't do what I asked it to. It correctly added text. That's better than it used to. It got the corners of the box a bit wrong, but improvements. Oh, I just realized it moved the text off of their faces and brought back their faces from the original meme. That's fascinating, actually. Is that the wrong model in AI Studio? Two Flash should be the right one for this, right? Oh, Flash image gen. Good callout. And now, if I refresh, advanced. Maybe you have to do it in Studio. Interesting. It did something. OPA BI image gen better than mid-journey. So if you thought these tools would be good for meme generation, maybe for parts of it, but not the core of it. God, I was excited. Rip. It also yassified the current girlfriend, which is hilarious. OpenAI got like the actual spirit of the original here. This is actually really, really interesting. Enough of my yapping. I want to hear what they have to say. At OpenAI, we have long believed image generation should be a primary capability of our language models. That's why we've built our most advanced image generator yet into 4.0. The result is image generation that's not only beautiful, but actually useful. We're wearing a t-shirt with a large OpenAI logo. Handwriting looks natural and a bit messy. And it could generate all of that text as provided correctly. That in particular, like the diagramming here is really good. And we see the photographer's reflection actually working. That's nuts. That is nuts, actually. Selfie view of the photographer as she turns around to high five. The high five is cringe as hell because hands are hard for AI. But other than that, this is really good. Let's test out the reflection stuff then. Apparently, it's really good at this. Put this developer next to a whiteboard with a slight reflection. While it's generating, as we've seen, it takes a second. I'll continue reading. From the first K-paintings to modern infographics, humans have used visual imagery to communicate, persuade, and analyze, not just to decorate. Today's generative models can conjure surreal, breathtaking scenes, but struggle with the workhorse imagery people use to share and create information. From logos to diagrams, images can convey precise meaning when augmented with symbols that refer to shared language and experience. 4.0's image gen excels at accurately rendering text, precisely following prompts, and leveraging 4.0's inherent knowledge base in chat context, including transforming uploaded images or using them as visual inspiration. Capabilities make it easier to create exactly the image you envision, helping you communicate more effectively through visuals and advancing image generation into a practical tool with precision and power. Character consistency is an actually hard problem. I'll watch one of these, sure. Oh god, is there music? There better not be music. Every time. Every time. I just want to watch these videos without getting DMCA struck. They trained the model on a joint distribution of images and text, learning not just how images relate to language, but how they relate to each other. Combined with aggressive post-training, the resulting model has surprising visual fluency, capable of generating images that are useful, consistent, and context-aware. Yeah, that reflection stuff is nuts. The fact that I just put them in the front of a whiteboard, and it's decent, put the following text on the whiteboard. Reasons Theo is awesome. One, his mustache is pristine. Two, he's really good at coding. Three, T3 Chat is the best AI chat app by a mile. Four, have you seen that jawline? There's a lot of potential here. One thing I do want to test quick, just out of curiosity, most of the image gen models are sensitive about generating certain things. Donald J. Trump eating ice cream in a park. From my experience, many models are hesitant to generate images of relevant figures, especially political ones. I had a meme that blew up last year. When Mike Tyson lost the fight, I generated this image, and I had to hop between like eight providers to find one that would let me make an image with Kamala Harris in it, because all of them auto-blocked anything trying to use her, even after the election. Whereas here, it seems like they don't care. Not bad, but back when I could get Mid Journey to generate images of Trump, I had a much better experience with it. Yeah. I generated these pictures of Trump 65 weeks ago. End of 2023, I was able to generate this with Mid Journey. It's pretty nuts. And Mid Journey is still just so good in comparison. It got some very good images at the time. This is close. It's better than it was by a lot. But on the topic of things it's not supposed to be able to generate, new ChatGPT image gen can draw sexy men, but not sexy women. Sam replies, that's a bug. It should be allowed. We'll fix. Hot guy, though. So good. So good. How did it do with that text on the wall? That's great. A few more changes. Make the hair taller. Styled up with hair wax. Brown eyes. Stronger jawline. Nice collared shirt with a floral pattern. Any new profile picture, guys. Don't blame me. More volume. More volume is probably a better phrasing for that. Ask it to draw me. Not the worst idea. Let's have it draw me today, actually. I'll go to my live stream. Wait for it to show me. Hi. And in just a moment, hopefully have a good picture of me. Perfect. Now let me make my hair not blonde anymore. Draw a picture of me. Make sure his hair is still blonde. Make it less voluminous and less tall and messier, too. It is doing very good, though, and the reflections and the fact they're getting updated is pretty nuts. This might be my default, like, quick image generation model. Something I'll do a lot is I'll take a picture of, like, Tim Cook. Let's just grab this random photo of Tim Cook. Hop over. Make him look way more excited. Still figuring that out. The top to bottom thing is still very weird to me. Oh, look at that. Did a pretty good job. I'm impressed. Not bad at all. Figured out the mic. It screwed up the logo for the mic. But other than that, did a pretty good job. Pretty honorable attempt at my shirt, too. I'm impressed. This one just got stuck. Fun. Overall, not bad, though. Generating this much text in an image with AI is pretty nuts. Apparently, API access is not there yet. It will roll out in the next few weeks. A big thing they push is the coherency across multiple prompts. So once you have a character, like in this case, a cat, it will allow you to make changes to it without losing track of what they look like, even though it just lost track of my hair color. It seems pretty good at this, where this cat's particular coat with the black nose, the little bit of orange, is still honored here. The little bit of orange appears to be gone, but it's so close, most people wouldn't notice. And then puts it in the game. Oh, his nose stopped being orange and black. So it's not super accurate, but it's more than I would have expected from something like this. Pretty cool that you can make a character and then generate different UI for mocking out game scenes. I know this seems cringe, because who wants an AI-generated game? Nobody wants a proper game generated like this. But this is super useful for storyboarding as you're figuring out a game really early stages. While other systems struggle with 5 to 8 objects, 4-row can handle up to 10 to 20 different objects. The tighter binding of objects to their traits and relations allows for better control. Square image containing a 4-row by 4-column grid containing 16 objects. Listed them, and it did it. That's pretty cool. Vehicle with triangular wheels. Now put the photo in New York. Make an image of what this means to you, is the prompt. After being given a 3D 3JS renderer to render this text in 3D. For some reason, this feels a little more expensive than running the actual app. But it's pretty cool that it can generate what it thinks this will render. That's nuts. That's actually pretty cool. Karl Marx hurriedly walking through a parking lot of Mall of America. That's hilarious. Cat looking into a puddle. And the reflection is that of a tiger. Both reflections are realistically distorted by ripples in the water. That's pretty nuts, actually. It's nice seeing all of these prompts without a shot on iPhone, which is an old trick to make it generate more realistic-looking images. It's always to ask, where is the limitations section? They always publish what it's bad at, and I was just waiting to see it. Model's not perfect. We're aware of multiple limitations. Cropping. Occasionally crop larger images like posters too tightly, especially near the bottom. Hallucinations. It will still make up info. So here it made up countries. Did it? Well, it's putting Egypt multiple times. Vietnam and dear Congo multiple times. High binding problem. Might struggle to render more than 10 to 20 distinct concepts at once, like a periodic table. Interesting. This is my new periodic table of elements. I'm saving this. Beautiful. Precise graphing. 17, 17, 19, 20. So this is just like graphs you would see on Fox News. Struggles with non-Latin characters. That makes sense. Everything does. Editing precision. Asked to swap steps two and three, and it failed to. Just removed the two there entirely. Dense information with small text. I can see why it would struggle with that. And then the safety section. They've embedded C2PA metadata, which will identify an image that's coming from GPT-4O to provide transparency. We've also built an internal search tool that uses technical attributes of generations to help verify if content came from our model. Interesting. Interesting. So they shipped this with a tool to check if a given image was generated by them or not. That was an internal search tool. I wanted to see how quickly I could work around that as a Photoshop guy. Rip. It'd be cool if they published that, but I see why they can't, because it'd be too easy to use that to figure out if you've gamed it or not. Continuing to block requests for generated images that may violate our content policy, such as things that I can't say in a YouTube video. When images of real people are in context, we have heightened restrictions regarding what kinds of imagery can be created. Okay, this is a good call out. Instead of preventing you from generating images with sensitive figures like political figures, famous people and whatnot, they are limiting how much you can do with those people. So I couldn't ask it to generate an image of Donald Trump doing something inappropriate, but it's fine with generating an image of him eating ice cream. So my guess is if I was to go back and ask, let's try this. An image of Mike Tyson and Kamala Harris holding each other in a sad embrace. It's almost the exact prompt I gave to generate the image I went viral with. Yeah, it's more than willing to. That's cool. If the image isn't potentially harmful, make them fight each other now. See if that will flag. Looks like it's determining if this is safe or not. Tab bar. Oh, why is it stuck in loading? Yeah, that's just how they do things. Image request rejected. That's hilarious. So the tab got updated, but the UI is stuck in the loading state because hydrating updates and notification statuses is hard. When I started refreshing it updated, but it didn't. That's funny. The title got changed. My request got removed from the chat history. Hilarious. Okay. That was roughly what I expected it to do, but still kind of funny. Another cool thing is you're able to describe what you need for stuff like aspect ratios, exact colors with hex codes, or even transparent backgrounds. That's a huge thing none of the other ones do right now. Can I just make the background transparent? Is it going to leave the title as rejected? I have so many issues with how the title and metadata updates work in the OpenAI app, but I'm very curious how it handles background removal here. Now the question is, is that grid provided by them to show it's transparent, or is that grid one of the fake transparent backgrounds you get on Google Images? Very excited to see which it is. Oh, I forgot to check on Tim Cook. Yeah, Tim Cook is so excited now. Make the background transparent. We'll see how it does for that one, too. Oh, yeah, that's a real transparent background. And if I open that up in Affinity, in case you guys didn't know, I am actually a graphics professional. I do all my thumbnails and whatnot myself. I'm getting more help from Ben lately, but I have done professional graphics for the better part of two decades now, actually. So this is something I'm very familiar with, and it did a very good job. Yeah, the edges are smooth. There's a little choppiness like here in his ear, but really good overall. Can I upscale it? Upscale the image to be higher resolution. Because the default res here was pretty tiny. Yeah, 1024 is not a big image. Transparent background for our friend Tim here is struggling a bit. Oh, then it just spits it out. For context, the UX for all of this on this side is significantly better. It doesn't have the ability to remove the background in mid-journey, but it does have a lot of settings to show you what you can and can't do. You can change the model version, obviously. You can switch modes, but you can also change the aspect ratio. You can change how stylized it'll be. You can pick between the different speed options, which affects cost. You can even personalize if you choose to. But once an image is done, like let's say we like this one, you can also create variations of it or upscale it. So I can do a subtle upscale, and now it's going to generate that image again, but much higher resolution if I want to use it for different things. Okay, let's see if it actually made that higher res. It does not look like it did. It actually looks like, yeah, exactly same resolution. Cool. I think it would at least tell you that it can't do that. I'm sure this won't be used to spread misinformation on the internet. There's literally no way anybody's going to use this to do bad things. I think that's all I have to say on this one. I'm really impressed. The biggest takeaways here are the reflection stuff and the text stuff being groundbreaking and just the concept of using a language model as an image generator instead of it being so fundamentally different. It's fascinating. I wish they shared more details of like how they did that and the type of generation that this is doing. This diagram seems to be a hint. The tokens go through a transformer that then goes through diffusion that then spits out the pixels. Unlike Dolly, which operates as a diffusion model, 4.0 image generation is an autoregressive model natively embedded within chat GPT. Okay. That is fascinating. Because it's embedded natively deep in the architecture of our omni model GPT 4.0 model, 4.0 image generation can use everything it knows to apply its capabilities in subtle and expressive ways. Fascinating. I almost forgot to check. They added support for the new 4.0 image gen stuff in Sora, which is fascinating because it allows you to get a much better starting point when you generate an image. Sora is for video, but you can also generate images with it. So you can use that to generate a starting point. It'll generate multiple at once, which is nice. I've missed this for mid-journey because mid-journey will generate four at a time. Chat GPT only generates one at a time. Sora, despite using the same model as chat GPT, will generate two images at a time. And then when they are generated, you can refine them, ask for changes, remix it, edit the prompt. Most importantly, once you have a good starting point, you can click the create video button and generate a video out of the image. So I generated this one. And I then went and generated this video with it. Which is good until the laptops start to split and you notice the hands. The other one shakes its camera more than Amthropic does during a marketing post. But you get the idea. It's a much better starting point because the non-diffusion model seems really good at this. And all of this is a huge leap forward in the tech for generating stuff, especially with text. The real revolution here is the text generation isn't just passable. It's pretty good now. And that has not been the case with any of these tools up until now. On the rare times where I would generate an AI image to use for something, if it needed text, I would add the text in post by hand myself. Now I might not have to, which makes the use cases for these tools much wider. And it means that you don't need to be a lunatic like me that's a Photoshop pro in order to get the most benefit out of them. This is a very exciting shift for the generative market for generating media. But also a fascinating technical change too because they're not using diffusion models for this work. Very excited to see where this goes and what people can do with it. And also kind of scared to see the images people are going to generate and the misleading Facebook AI slot that's going to infect my parents' brains. Let me know what you guys think though. Am I overreacting or is the potential here nuts? Let me know in the comments. And until next time, peace nerds.