Open transcript
The OpenAI servers are melting. This is due in part to their newest addition to their GBT4O model. 4O is a multi-modal model, meaning that it's trained on not just text, but also visuals, sound, speech, etc. The O stands for Omni, like all, as in the one model to rule them all. And the latest feature allows you to turn everything into a GBT Studios style image. Actually, that's not all that it allows it to do. You can do anything with it, like anything, but for some reason, I don't know why. I understand it. I just don't know why this exactly, but everyone is making a Ghibli style images of themselves, of their pets, of various family members, of any event you can think of that had a picture of it is now a Ghibli style image. By the way, I would also argue that Ghibli is an acceptable pronunciation. Maybe we might get to that later. But the demand for this new feature has been so high that they're not quite releasing it to the free users yet. You have to be on one of the paid plans to participate. And the demand has been so overwhelming that they're not releasing it to the free quite yet. You have to wait just a little bit. CNBC posted an article called, Now, unfortunately, this was not without consequences. This person on Twitter just said, I got this cease and desist from Studio Ghibli. AI creators deserve protection, not punishment. Expression is sacred. Imagination is not illegal. If I have to be a martyr to prove that, so be it. I'm assembling an illegal team. Firms who believe in this fight reach out. And here's a letter from Studio Ghibli. As you can see here, it's unauthorized use of Studio Ghibli intellectual property. Cease and desist. If you're not aware, Studio Ghibli is behind such great hits such as Spirited Away, My Neighbor Totoro, Kiki's Delivery Service, That's Damn. They have a lot of good stuff. Very well-known, very respected anime studio. Lots of people attack him saying he basically deserves it. He's stealing art. He's looking to profit off of somebody's art. Right? Saying it's deserved for stealing art. There's a lot of anger and aggression directed at this person, even though he's on the receiving end of this letter. As one of the people said, their letter is designed so much better than your app. And the person replies, Funny you mentioned that. So on a closer look, this is not a real letter. This is not a real company. This is not a real phone number. It's not a real URL or email or none of this is real. The letter is, as far as you can tell, probably generated with chat GPT. Or at least that would be the most entertaining of all things. Either that or just, you know, typed up by the guy. Either way, exceptional, hilarious troll. Very, very good. But everybody's very impressed with this new image generation within chat GPT. Here's the founder of Shopify. Impressed with this. I tried to see if it could recreate that severance idea of somebody working on a computer inside their own head. I feel like it did pretty good. Here I have to create a Wikipedia page about recursion with a screenshot of that Wikipedia page as a sort of the main image. So as you can see, it creates that kind of infinite effect. And as far as I can tell, like if this is the quote unquote page, everything is perfect. As far as I can tell, everything is written out. At least I'm saying in terms of it's not making any obvious misspellings or anything like that. Maybe I'm missing some. But if we zoom in, notice if this is kind of like the second iteration of the page, things start getting a little bit more glitchy, right? So it still gets computer science, search, search, Wikipedia, recursion. So there's a lot of things that are right. But then more and more it starts breaking down. Of course, if you go even deeper, I don't know if I can even zoom in that much. Yeah. So this here, that's like the third, it's the third layer. So here it starts kind of breaking down towards blurry, you still see recursion, etc. And then continues. Like I say, this is, this is pretty impressive, right? So this is on the left here, this is the original page, right? And then it goes kind of deep into multiple layers. Pretty cool. Someone recreated this image in the anime style. So I'm not sure if this is Ghibli or not, but this is definitely sort of that kind of cutesy anime style. Carlos Perez and the GPT-40 meme generation. Very, very good. It said that when it got released, all the other AI image generation companies, they came out and said, I declare bankruptcy. Not really, but I just like this meme. By the way, just making it into a Ghibli style image, does that automatically make it a meme? Interestingly, Pliny, the liberator, made it create a whiteboard with the solved Riemann hypothesis. So there it is. After all these years, it has finally been solved. Yep. And I bet it's right too. This was the image for the Nobel Prize in Physics 2024. Somebody thought they left the person out. So they went ahead and fixed it for them. There's Jürgen Schmidhuber. This is great. It really captures the style because this was drawn, I guess, for the Nobel Prize for the ceremony. This was drawn in whatever style that you would call that. I feel like it nails the style. And I mean, honestly, if I was being honest, I'd say it probably kind of improved on it. I pulled up the original image in a higher resolution. I hate the fact that it has the person that illustrated on it. I mean, it's great. I love it. But I feel like this is the same sort of art style, but the features are more true to life. By the way, sometimes people get upset if somebody says that some AI output is good or bad. I think a lot of people in the artist community think that if you say this piece of AI-generated art is good, there's a little bit of like hostility towards that, saying that, well, you shouldn't be pro-AI. This is not pro or against AI. I'm just saying that we're getting the point where it might start exceeding a human artist's abilities. I think that's just a factual statement. Whether you think it's a good thing or a bad thing, I think most people are going to be agreeing that sort of the technical abilities of it are going to get better. Logan Kilpatrick of Google was asking, how's everybody doing what they're using GPT 2.5 Pro today? Sam Altman responded, saying, my whole timeline is very impressive images. Great ship, of course, kind of being funny because it's OpenAI that shipped the product that's flooding everybody's timelines. And of course, someone immediately makes that into a Ghibli-style screenshot. Like I said, everything's getting Giblified. There's no stopping this wave. The reason for that is simple. You have a choice between, you know, getting some actual work done or Giblify everything. Of course, you know which road we all chose. If you've ever heard me say please and thank you to the models as I'm asking to do stuff for me and testing them, I guess this is sort of the image that I have in my head, maybe of some not too far distant future, where as my life is flashing before my eyes, one of the robots puts out his hand and says, spare this one. He was always polite to Chad GBT, saying please and thank you. It can't hurt to be polite, especially if you, you know, if you have any concerns about the future. This one was interesting. So this is a view of Ringworld. And this is something that I've never been able to do with any model. It's interesting because, I mean, I wouldn't rate this as a good image. But you've got to understand, this is a one-of-a-kind image. Because it's the only one that I've seen generated that looks like a Ringworld. Ringworld, if you're not familiar, imagine a sun in the center of the galaxy. But instead of planets around it, you have this ring, right, kind of at the same sort of distance from the sun as, let's say, the Earth would be, right? And so all the inhabitants of that ringworld live on the inner part of it. So the sun is shining on them. And as you can imagine, a world like that would be just gargantuan, like insane proportions. I tried to recreate this. I was not able to. Here's kind of the closest that I got. So again, not quite what we're looking for. It's just not really something these models can produce. We have phenomenal examples of various events, memes, etc. in Ghibli style. Incredibly, it can do a depth map, right? So if you give it, I mean, you can tell what image this is, right? But instead of changing it somehow, we're able to create a depth map that kind of tells you what's in the foreground, what's in the background. I guess this shouldn't be that impressive, but it just is. If you watch the video where we cover that paper, that AI paper out of Harvard called Beyond Surface Statistics, this is one of the really mind-bending and weird things to understand about AI and neural nets. In our experiment, they fed the image generation model just a ton of images, but they were all 2D images, right? So they didn't have any depth data, right? It was just 2D images, right? Pixels on a screen, basically, right? And it was combined with words. So it's like pictures of cars that said, this is a car, red car, blue car, whatever. So it's very similar to how we train other image generation models. But what they found is over time, as it got better and it was producing coherent images, when they tried to see kind of like cut into its brain and figure out how it was thinking about building these things, I believe they used what's called a linear probe to do that. But what they found is that that model developed some ideas, some mental models, if you will, about depth, about what the sort of the main object in that image was. So they do sort of learn to implicitly understand the world. If they need to create images, they need to kind of understand or have some grasp of what the 3D space looks like, of what light and shadows are, etc. So they do develop a sort of like some representation that allows them to do that. So I don't know how OpenAI trained this one. Maybe there was a depth information that allowed it to do it. But more likely than not, it just figured it out sort of implicitly. In other words, we didn't teach it how to do that. It sort of just learned it on its own, which is kind of interesting to think about. And of course, here's that same meme in a lot of different formats. Looks like we got the strings, puppeteer dolls on there, old school Disney animation, etc. This one, it actually turns the person all the way around, which is similar to how they used to draw things back in the days. You wouldn't have people like at weird angles or looking over the shoulder. They were kind of facing one way. So this seems like a great representation of what that would look like. Make an entire iMessage sticker set of yourself with 4.0, right? Turn me into a chibi sticker set, right? So now you're able to text with your friends, with your own customized emojis and whatnot. I mean, that's kind of cool. Also, it seemingly creates great coherent multi-frame sprites, right? So if you're able to create like four sprites of some image, then you're able to sort of animate that image like these spinning coins here. That means that it can be dropped right into a game engine. So this can generate actual video game assets on the fly for you, it seems like. Tons of people have generated various UI elements. You saw my thing with the Wikipedia, but there's many, many more, including apps, websites, you name it. You can take a model, give them a drink and say, you know, have the model holding a drink. And it tends to work very, very well. This is a reverse Ghibli sort of thing, making those characters look lifelike. Totoro in real life would be a nightmare fuel, I feel like. Actually, the cat bus would be the real nightmare fuel. Imagine me, that thing and a dark night in the forest. That's going to turn you off to cats for the rest of your life for sure. And the other thing that is kind of cool is using the various AI video platforms to take those images and turn them into videos. This is an example of, you know, video creation for a podcast, which I got to say would be pretty fun to watch, I feel like. Here's that same scene from the Severance TV show. By the way, if you're not watching it, why are you not watching it? It's pretty dang good. I tend to kind of hate most of the stuff that comes out now. I don't know. I guess you get old, you start hating new things. Who knows? But Severance, Severance is on point. They really hit on something that I can't quite put my finger on it. I just started season two, so no spoilers. I'm serious. And of course, Lord of the Rings, Ghibli style. So somebody created a full-blown, this is a few minutes long, sort of Ghibli style animation of Lord of the Rings. And it gets quite involved, right? Somewhere in there is Balrog, right? You shall not pass. Very, very good. I'd be horrible at dubbing. I was so off on that one. But you get it. For the people watching it with the AI dubbing the languages, that wasn't the AI that failed. That was me. I was off. How they dubbed Squid Game so well, I will never understand that. Now, of course, not everyone is happy about this trend. As you can see, there's a lot of people that are kind of unhappy about it. And again, I am not here to change anyone's opinions. I understand that people can have negative feelings towards AI-generated art or images like that. My stance has always been that we have to sort of, as a society, kind of figure out where we stand in relation to this, right? How do we treat AI text versus human-written text? How do we generate AI music versus human-written music, et cetera? The only thing that I kind of disagree with is the idea that we're breaking copyright by training the models on, you know, whatever text or images that it looks at. Reading of something was never a copyright violation, right? Taking a book and doing a scan of it, you know, making a photocopy, that was not a copyright violation. Like, trying to sell that for money, that's a copyright violation. You know, if you have a website up, there's tons of crawlers that go through it and read, quote, unquote, all the text, whether that's Google bots, Bing bots, or whatever other bots are out there that are crawling constantly. No one's getting sued for copyright violations. We've sort of determined that the machines are able to comb through data. We don't consider that a breach of copyright. Copyright occurs when we reproduce stuff, all right? Either for commercial gain or in a way that intrudes on the original copyright holder. So I'm sure some of the data that this model was trained on was, you know, Studio Ghibli's art. I personally don't see that it's a violation of copyright. But now that a million people, at least all over the world, are reproducing that art using simple prompts, you know, that's where the discussion lies, I feel like. What is and isn't okay, I think whatever side of the issue you're on, I think most people are going to agree that maybe our current existing copyright laws are maybe a little bit outdated. So interestingly, Japan, I believe, came out and explicitly said that you're able to train your models on whatever data and it's not considered copyright infringement. Currently, OpenAI is telling sort of the U.S. government, kind of saying the same thing, that in order to be competitive, we have to be able to train on whatever publicly available data there is. So I doubt that we're going to see any sort of resistance to that idea. But now the outputs of the model, you know, if they're one-shotting somebody else's book, somebody else's copyright, materials, etc., that's really where I think the discussion is. And that's where we really need more clear laws. Here's Satya Nadella, CEO of Microsoft, going, look, all I know is I'm good for my 80 billion. So if I used a chat GPT to create something like this, who holds the copyrights? Do I? Does Satya or perhaps Microsoft? Is it OpenAI? Is it Studio Ghibli? Because I think you might be hard-pressed to describe exactly what about this makes it their copyright. Now, of course, if I used that word when I was describing it in the prompt, then you have sort of more data points. But if you don't know where it came from, it's kind of hard to assume that this was done referencing one artist or another. But interestingly, Hayao Miyazaki, who was the founder of Studio Ghibli, he did have an opinion about machine-created art. By the way, Ghibli was an Italian aircraft. That was the nickname of the aircraft, a name for a Libyan desert wind. It sounds like the founder of Ghibli Studios. That's what he was naming his company after. It was kind of like a reference to it. It's just the pronunciation was a soft J, so it was Ghibli. And I think the guy that was one of the creators of the airplane, he's actually featured in some of the movies. So I don't know. I'd make a case that you can say Ghibli or Ghibli. I think both are somewhat correct. But again, GIF, GIF, I don't care. But here's what the founder of Ghibli Studios said. Ghibli was so much less experienced than the technology. It was complicated. And then... The fact is true. I think he was pronunciation when he smelled. Not really? That's what he decided for. I'vezk Jadiwa INice-aheen-ahe are sunlight-hai. Ghibli-gasaw, I-ahe is called the Flash-tiene missions. When Iingen is called Matt Badre, I-ahe here, I'm like... I can make more relief from you to land on Italy. Yikes, kind of feel bad for this guy. Although it would be kind of funny if he's now working for some AI company trying to create the greatest AI image generation machine of all times. Now I think it's important to understand that that statement took place in 2016 before they were generative art models, before there were the Transformer or anything like that. Someone Ghibli'd Mike Tyson. I wonder who it was. Oh, Mike Tyson. I mean, you know his stuff is going viral when like everybody's posting about it. So now there's a lot of articles talking about it, talking about copyright laws. Some people are saying that the Ghibli studios should sue OpenAI and whoever is making these images. Interestingly, there might be another side to this thing. And you might be thinking that OpenAI kind of goofed up by doing this. And maybe that's the case. But as Grant Slatton said, honestly, OpenAI is incredibly fortunate. The Positive Vibes or Ghibli was the first viral use of their model. And that's some awful deepfake nonsense. So it's important to understand that this model is a little bit more unhinged, so to speak. So they are trying a new approach. There's a person on the research staff of OpenAI. She posted a whole blog post of kind of their new approach to doing guardrails and AI safety. And no longer are they doing sort of these blanketed, broad sweeping denials to certain prompts. And in fact, there's been people that have generated stuff that I can't even show you on this channel. And on the OpenAI live stream, they did mention that. They were saying that, yeah, you're going to be able to push this model to produce stuff that was before would have been denied. And if you're on Twitter, you probably saw some of that stuff. It will definitely do more than the other models unless you go to kind of the open source, like models that are specifically made for the not safe for work content. So I'm not sure how it compares to like Grok, for example. But because I haven't sat there and just put naughty things in there and see which one comes up with it or not, I just haven't done that. I'm sure somebody will and I will report on it. But definitely this has an ability to produce a lot of deepfakes and a lot of stuff that probably could have been spun out of control, whether it was like inappropriate or deepfake or something that tricked the people into believing something that was false. But as I'm responding here, he says, believe it or not, we put quite a lot of thought into the initial examples we show when we introduce new technology. Their big thing, their big demo, if you recall, was using it to create an anime image from a selfie that they just took. It was very impressive and a lot of people decided to replicate that. That became the trend, the meme, the whole world is getting gibbified. But the point is, since this is the trending story, this is what everybody's doing, whatever he's talking about, it really kind of does put a damper on all the bad stuff that could potentially happen. These are very positive vibes. Even when people take horrific scenes and kind of gibbify them, you kind of look at them and go, ah, it's not that bad. It seems to just kind of tone everything down. At this point, pretty much most of the people in the AI space have gibbified themselves. Everybody's looking good at this thing, except me. I have not been able to generate anything that remotely looks like me. I don't look like this. I don't look like this. I don't look like this. I don't look like this. I don't look like this. And why am I angry? I wasn't angry in the picture. This one's kind of cool. And it turned my profile picture into, which I guess is probably one of the better ones. So I'll probably stick with this one. This one's pretty cool too. It noticed I had the gold chain. Just a nice little attention to detail. I guess it could have been worse. It might have made me look like Totoro. But what do you think? Are you jumping on this latest trend? Are you enjoying it? Did you gibbify yourself? Are you overall impressed with the new image generation abilities? Let me know in the comments. If you disagree with how these AI tools can generate images like that and use somebody else's style, specifically referring to that person's style, where do you think the problem lies? Is it the people that are using it? Is it the fact that the AI tool can actually do those outputs? Or do you look even sort of deeper, more fundamentally, saying that these models shouldn't be able to train on other people's copyright works? Let me know in the comments. If you made it this far, thank you so much for watching. My name is Wes Roth and I'll see you next time.