Open transcript
OpenAI just finally dropped O3 Pro. And as excited as I am about that, it's my second favorite thing they did today, because they also cut the price of O3, the standard one, not O3 Mini, by 80%, making it cheaper than GPT-40, Claude, all of the Claude models really, and even Gemini 2.5 Pro. Considering how smart O3 is, that's insane. But O3 Pro goes even further. It's not anywhere near as cheap, but it is 87% cheaper than the previous O1 Pro model, which is just absurd. And the quality of the responses you get from it is insane too. As smart as O3 Pro is, it tends to take a lot of time reasoning, like almost four minutes on a simple high prompt. That's all tokens that we have to eat in our testing, and someone's got to foot the bill. So we're going to do a quick word from today's sponsor, and then we'll dive right in. AI is really good at parsing data from structured formats like YAML and JSON. It's a lot less good at getting data out of webpages, especially if they're client rendered. Turning complex HTML into something that your AI can use is not an easy task, unless you use today's sponsor, Firecrawl. These guys make it way too easy. You give them a URL, and you get back JSON, or whatever other format you might want to. It is hilariously easy to get started. They even built a nice TypeScript-friendly SDK. And when I say TypeScript-friendly, I mean it. You might be questioning, wait, how does it know what format the data is going to be when you get it back? That is one of the coolest parts. When you use it in JSON mode, you can pass it a ZOD schema. And now when you call the scrape function with that schema, you get back data in the format that you ask for. That's so cool. And the use cases I'm coming up with just looking at it are absurd. I love the is NYC because we might actually start using this to figure out things that we want to invest in. It's smart enough to wait until content is loaded and rendered. So if you're waiting for the JavaScript to run, they got you. And it's not even that expensive. When I saw it, I immediately assumed it was going to be massively overpriced. But for 16 bucks, you get 3,000 scrapes. And the free plan includes 500 credits. Each one is good for a full page. That's crazy. I just keep coming up with use cases where this is valuable, and I'm sure you will as well. Check them out today at soydev.link slash firecrawl. I hope you'll never complain about ad breaks when they are shorter than the reasoning time for some of these models. But the reason it takes so long is it's thinking a lot, even when it's possible it shouldn't be like it did here. Let's dive into the release notes of O3 Pro and then go into O3 because I'm really excited about the price change. It kind of changes the whole landscape. But first, O3 Pro. It's now available for Pro users in ChatGPT as well as in their API. They have Plus for 20 bucks a month and Pro for 200 bucks a month. Fine. But a quick reminder, eight bucks a month for T3 Chat. And you now get access to O3 as part of the core plan for that eight bucks a month. If you want to try it out for cheaper, use code O3, please. O3-PLS at checkout for just $1 a month. It'll only work if you're not already a subscriber. Anyways, let's talk about these models because things are really, really cool. Like O1 Pro, O3 Pro is a version of our most intelligent model, O3, designed to think longer and provide the most reliable responses. It's been a bit since we saw the best in class, so to speak, from OpenAI. Their reasoning models are obviously their best models, but they're broken into three tiers effectively. Right now, when a new base is made for their reasoning models, they usually ship it under the mini tier initially. They have mini, they have main, which is just the term I'm going to use for the one in between, and they have pro. These are the three different tiers. The difference is effectively the size of the model itself, the amount of tokens that it's referencing when it generates. It's a lot of different things go into it. It's also how much time it's going to be spent reasoning, how much GPU time and GPU energy it costs to actually run the model because it's traversing a larger set of data. But think of this like with Llama 4, how they have the Maverick and Scout versions, or how with other models they have the small, like the 24 billion parameters and the 200 billion parameter version. They split this into three sizes, which means they have different characteristics. Since the mini versions are smaller, they tend to be much faster and also much cheaper to run. The main ones tend to be a solid balance, and then the pro ones tend to be a good bit smarter, but way slower and more expensive to get things done with. I found so many bugs in the ChatGPT site using 01 Pro because it was so slow, it would regularly get the site into a bad state. And I would regularly have problems where it just failed to finish the generation step because I went to a different tab or started a second chat while that one was still going. It was not a pleasant experience. And thankfully, they've made a lot of improvements. Shout out to Naman. But yeah, the pro models are, funny enough, bad user experience. The intuition for most people is, oh, I just want to go use the smartest one. I'm going to go down the chain and use whatever I can for my problems. The further you go down this, the worse experience users have, which is kind of crazy if you think about it. The smarter models are worse UX. But if you have a problem that they can solve and the other ones can't, it can benefit you a lot. What doesn't make as much sense to me is why 4.0 is still on top before the rest of these, because 4.0 is a lot dumber than these other models. But it also, because it doesn't have reasoning, it starts to generate a response faster. And they've put a lot of effort into fine tuning how personal it feels. 4.0, since it's not reasoning, it's just immediately responding, feels better to have a back and forth conversation with. And it's not overthinking your responses. I find all of the reasoning models from OpenAI to be much more clinical, which I personally massively prefer. I don't want my AI to sound really human, personally speaking, but I'm not using it as like a friend I talk to. I have too many people I already owe a text. But for most general people, like my mom would probably like 4.0 a lot more than the other models. And she's not asking anything that you need a smarter model for. And for things like voice chat, having that immediate response without reasoning is really, really nice. That said, the mini models have gotten so good and so fast, that it's hard for me to not default to 03 and 04 mini nowadays. So this gives us a picture of where things are at. But there's also the horizontal spectrum of like 01, they skipped 02, because they didn't want to get sued by a particular phone carrier. And they did 03, and they did 04. What's interesting is, for various reasons, they tend to put out the mini model first now, which means for a little bit, we had 01 Pro, and just like, you know what, I'll find the real numbers from their benchmarks. Why not? So here are benchmarks across the different versions of 03, 01. I also left in Gemini 2.5 and Claude Force Sonnet here. And you'll see for the basic like math benchmarks and AIME, the consistent improvement. 01 mini was slightly worse than 01. 03 mini was slightly worse than 04 mini. And 03 beats out most of these. In the more general artificial analysis index, you'll see it's a pretty steady improvement. The thing that's interesting is that 03 came out for like general use after 04 mini. And 03 Pro just came out today, even though we've had 04 mini for a while. And they're pretty close in capabilities. The thing that makes this feel weird is that these higher tiers take so much longer to come out, that by the time 03 base or Pro comes out, 04 mini already exists. And the level up on this axis, sometimes matches or even cancels out the level up on this axis. It just kind of feels weird that 03 Pro happens today, when 04 mini already exists and is such a good model. 04 mini is in my default in T3 chat for a bit. It used to be 2.5 flash, I still use it here and there. But more and more, 04 mini is fast enough and so consistently good that I find myself reaching for it by default. I'm also using 4.1 in my editor. I'll talk about that some other time. I have some fun tips there. But for now, we're focusing on these reasoning models. And 03 Pro seems to be groundbreakingly smart for specific things. But it does just feel weird to have 04 mini that is so much cheaper than 03 Pro and feels just as smart for general usage, and also is way faster. That said, seems like OpenAI is kind of thinking of these as entirely different use cases. Rather than seeing this as small, medium, large, OpenAI sees this as general, more deep dive, deep research type things, and literally writes reports for you. 03 Pro has a 64% win rate versus 03 on human testers. That doesn't mean it is 64% better. That means it beats 03 64% of the time, which means for most people, most of the time, it's close. For the win rate to only be a 14 points higher is kind of crazy. But it shows how good all these models have gotten. It'd be really interesting to see 04 mini in comparison there. OpenAI has been really good about showing how to think about prompting with these new models because it is quite different. Even stuff like your system prompt, they officially recommend that you add as little system prompt as possible in order to let the context be focused on the actual problem that you're trying to solve. So here is a good example of the anatomy. You start with a goal, you get specific about return format, you give any warnings or details you think are important, you dump all the context that it needs, and then you give it capabilities like MCPs, tool calls, and whatever else it might need to do. And when you do things right, the benchmarks are pretty crazy. As we see here, scores a decent bit higher with human testers versus 03. But also the traditional benchmarks like competitive math and whatnot, it's beating out these other very, very smart models. 01 Pro just kind of feels like a weird moment in time now because it was so expensive. And there are so many things that are better than it now. It's kind of crazy. Oh, this is funny. Temporary chats are disabled for 03 Pro as they resolve a technical issue. I know what that is. Since the temp chats aren't persisted, it relies entirely on streaming to the client when it's done. But the state of streaming to the client on these really slow generations is probably entirely broken. So it can't do that. That's really funny. It shows how hard it is to build these types of things. This is one of the things I'm thinking about for our temp chat solution too. So sympathy to the engineers who are stuck solving that problem. It's not going to be fun to do. It does also seem to be significantly better at code, which is exciting. I am very curious to see how it feels to use 03 Pro versus 03 solving problems inside of cursor. Should be fun. I personally like to wait to get more thorough benchmarks from folks like artificial analysis. But as previously mentioned, the reasoning times are insane. So yeah, this isn't going to be cheap to run and it's not going to be very fast either. All of this said, Gemini 25 Pro still beats it on some of these benchmarks, some of the math ones in particular, but they are neck and neck. It's interesting. Also, 04 Mini performed really well, and I'm not surprised they didn't put it in the benchmarks they're sharing here because 04 Mini probably wins in a handful of these things. 04 Pro is going to be a very interesting release when they get there. This was a fun interaction happened at the beginning of the year from the latent space guys where the author, Ben, was not sure about 01 and the O models from OpenAI initially. It got ratioed by Sam Altman and quote tweeted by GDB about it. And the lesson that Ben learned is that you should not use these models as a chat. You should treat it like a report generator. Give it context, give it a goal and let it rip. That's exactly how I use 03 today. But therein lies the problem with evaluating 03 Pro. It's smarter, much smarter. In order to see that, you need to give it a lot more context. And I'm running out of context. There was no simple test or question I could ask it that blew me away. We're well past the days of my favorite ball bouncing test. There's a reason I didn't open the video with that. It's just not really relevant anymore. It barely was then and it isn't now. But yeah. The co-founder Alexis and I took the time to assemble a history of all our past planning meetings, all of our goals, even recorded voice memos and asked 03 Pro to come up with a plan. We were blown away. It spit out the exact kind of concrete plan and analysis that I've always wanted an LLM to create, complete with target metrics, timelines, what to prioritize, and strict instructions on what to absolutely cut. The plan that 03 gave us was plausible and reasonable, but the plan 03 Pro gave us was specific and rooted enough that it actually changed how we're thinking about our future. This is hard to capture in an eval. I agree. I feel like the capabilities of these things are no longer truly traditionally benchmarkable. There's a reason that the people deep in that space, like my stupid safety bench so much. Snitch Bench, as silly as it is, is a weird practical edge case deep dive on what behaviors these models do and don't have. I feel like benches are going to move away from how smart is this model and towards how does this behave in different scenarios because that's what's going to matter more and more. And 03 Pro might be way smarter. That doesn't mean you should use it as your default model because it isn't good to chat with. Trying out 03 Pro made me realize that models today are so good in isolation that we're running out of simple tests. The real challenge is integrating them into society. It's almost like a really high IQ 12-year-old going to college. They might be smart, but they're not a useful employee if they can't integrate. That's a hilarious way of putting it. And as the author then says, this integration comes down primarily to tool calls. And historically, OpenAI's models have not been great at tools. That has since changed to the point where now when you go to ChatGPT, they have a little button here that is choose tool. And you can pick which tools you want to give access to on their website, which means we need to start adding more tool call options for T3Chat. More coming soon. It's interesting that deep research isn't considered a tool. It's just here separate. But I have 250 available during my time with my $200 subscription. So that's about a dollar per report. Fascinating. You can also link it to external sources. I've seen them leaning more and more into this. It feels like the tool calling stuff and the large context stuff went from out of OpenAI's interests to the thing they are most focused on. And they're going ham on it. It's still so crazy that there's a tool call button inside of the ChatGPT site. I never thought I would see the day. But it really highlights how they're thinking differently about things. Speaking of which, look at all of the model options that we have now. RIP 4.5 is getting sunset soon. Apparently, O3 Pro is better at discerning its environment, as in what it has access to and what it's supposed to be doing as a result. It can more meaningfully communicate what tools it has access to. This is one of the biggest problems I have with Gemini models, especially the recent ones. They love to describe what they think they're going to do with the tools and then just not do it. And this got way worse when they added reasoning summaries, because when they added that, it seems like they deprecated the deeper reasoning and cursor. So now it doesn't get to do as much as it used to. It's quite annoying. It also seems better about when to ask questions about the outside world instead of hallucinating information. That said, as these models have gotten smarter, they actually seem to hallucinate a bit more, which is scary. And I more and more think the solution to hallucinations is giving these models access to tools and the ability to stop and ask a question to collect additional context. There are problems with this, though, and they kind of come down to that pricing model. And it's the Edwins. Like, 03 is now cheaper than 40. It's also, 04 Pro is significantly cheaper than 01 Pro. Even though it is hilariously cheaper than it used to be, it's still quite a bit of money for those input and output tokens. That's $20 per million input and $80 per million out. So imagine you have 150,000 tokens of context. Let's model this out. We have this 150k token context. It starts with like the instructions. So this is what to do. And the rest is tons of context. So we give all of this to a model like 03. 03 reads through the context, looks at what we want it to do. And then it responds with a very short response, very short, costs almost nothing, asking, what about x, where it wants to know more info about this particular thing that it couldn't find in the context window. So you tell it, I'll use colors for this green is from them. And yellow is from us. So yellow is billed at 20 per mil. And green is billed at 80 per mil for 03 Pro. Just inputting this message is $3 of API costs. If you have enough context, this response is nothing. It's literally going to round to like zero, like, not even a fifth of a cent, because it's so little like information, the response, it barely matters. But if you now respond saying, Oh, x is whatever, even if this is really low context, we'll say this is like 1000 tokens that you've added here. Oh, that means it's gonna be really cheap, right? Sadly, wrong, because the next request has to have the whole history for all of the context. So instead of this being very simple, where this was like 400 tokens reply, and then you give it 1000 tokens, and then it generates a new response, it effectively has to reingest the whole thing. So as the history gets longer, this $3 gets rebuild on every further message. So it kind of sucks when a model stops and asks, because now it has to reingest all the context in order to add more stuff. This is part of what makes tool calls cool. Because if instead of this reprompting you, it halts itself, it goes out during this execution, get some more data, and then adds that back in as additional context, it doesn't have to lose and then rebuild the context that it had before. So tool calls are cool, but human in the loop has historically kind of sucked, because it forces this context to die and then be remade. This is also why caching is so important, because one of the solutions that's implemented in stuff like as of recent, they got this with Gemini models, this will automatically be cached in the cache reads cost less, but it's only cached for a certain amount of time. So if I wait 31 minutes before responding here, now I have to eat this ingest cost again. But if I do it in 29 minutes, it'll hit the cache and cost significantly less. But now you have to be very intelligent about when and how you cache, you have to think a lot about these things. And it just disincentivizes you from doing human in the loop on these smarter models, because every time you stop, you're risking eating this bill again for the whole context. This should make you sympathize with people like the cursor guys, because when you ask the model to do something slightly different, or make a small change, it still has to re ingest everything. So if you ask a model, generate this whole code base for me, and then you ask it, make this one change, the input cost is still going to be the same for both. Obviously, when you're generating that much, the output token costs starts to go up a ton. But from my experience, the input tokens have been very expensive. And for something like cursor, I would imagine those input tokens are the vast majority of your expenses. Finding the right way to manage your context is becoming more and more a problem as the use cases for these models become more powerful, but also need you to confirm things more often. That said, as the prices race down, these problems matter less and less. It's been kind of crazy watching the like what I call the race to the bottom in the pricing wars here. We hop over to the cost to performance here intelligence versus price. Suddenly, Claude for Opus looks really bad because it's so expensive. And it's not much smarter. It's actually quite a bit dumber than some of these other options. I had to turn off Claude for Opus to make this chart even like readable. This actually makes this very funny now because the Sonnet thinking model and Sonnet in general from Anthropic is their second most expensive model. And it's still the most expensive thing on this chart now, which is kind of crazy. I never thought I would see the day where everyone would undercut them. Okay, I kind of did, but I didn't think it would come so fast and so aggressively. I'm going to turn off more useless things in here. When you open this chart this way, O3's new price starts to make a lot more sense. They're really neck and neck with the most recent Gemini 2.5 Pro refresh. And I think the new pricing is meant to undercut them a bit. I do trust what they say, which is that they found ways to make the inference steps way cheaper and faster on their end. O3 is running faster. I saw some people speculating that they actually swapped out the model and there's a new dumber model that you're running instead. Otherwise, it wouldn't be so much faster. No, they made changes that make it faster to run. Those same changes also make it cheaper to run. Faster models tend to be cheaper overall because they're doing less inference. They're spending less time on that compute step, which is why these models cost so much money to run. That's also why 2.5 and 2.0 flash from Gemini are so fast because it runs absurdly quickly. Same with Grok 3 Mini. These models are cheap for a bunch of reasons, but the biggest one is that they're spending less time and less energy doing compute. This chart has gotten so interesting so quickly though. O4 Mini still stands out as absurdly intelligent and reasonably priced. That said, the costs here are not accurate to real world use because the amount of reasoning it does will greatly inflate the price. Some of these models might solve in 1,000 reasoning tokens what others do in 10,000 reasoning tokens, which makes the costs go up a ton accordingly. This chart's the one that really tells the story. This is how expensive it is to run the artificial analysis benchmark, which is how they measure these models. Running for Gemini 2.5 Pro ended up being the most expensive, even more so than Claude 4 Opus, because it just generated so many more tokens. The reasoning cost was $852, where with Claude 4 Opus, they don't split up the reasoning and non-reasoning cost. It's a bit harder to split, but it ended up being half as expensive almost overall. So even though Opus is more expensive per token, Gemini 2.5 Pro does more reasoning, so it's generating more output tokens, which makes it worse. So the traditional cost per token is no longer the only number you have to look at. You have to look at a lot of other things. If we compare running 2.5 Flash reasoning to 2.0 Flash on the same benchmark, remember, 2.0 Flash costs $3. 2.5 Flash reasoning costs $319. Massive gap. Massive gap. That said, once again, O3 Mini is coming in clutch, being a solid in-between. The fact that O3 Mini is cheaper than 2.5 Flash reasoning to run this benchmark is kind of hilarious, considering how much more expensive it is per token. But the per token numbers are starting to matter less and less, which is kind of weird. I'm also excited to see what this looks like when they rerun the O3 Bench with the new pricing, because I bet this will get pretty cheap. This actually might be the rerun. As they show in the official artificial analysis charts, O3 has shifted really far left in the price per million tokens. But this was the fresh run with the new pricing. It is hilariously cheaper. I also love seeing Quen 3 so high up here, because the actual token costs for Quen 3 are really cheap. But as we've established with my favorite weight index, the amount of times it gaslights itself, solving any given problem, results in Quen 3 just generating endless fucking tokens. So you end up eating a bunch of cost and waiting a bunch of time for the model to generate an answer. This is the most fun chart, though. This is intelligence relative to cost to run the benchmark. And funny enough, the only thing in this top left quadrant of good score relative to its cost to run is Grok 3 Mini again. I hate that Grok 3 Mini continues to be a weirdly good value solution, but it really does. O3 and Gemini 2.5 Pro are the top performing overall, but O3 is now quite a bit cheaper than the 2.5 Pro option. And funny enough, O3 and O4 Mini cost very similar amounts to run this benchmark on now, which I never thought I would see the day that it's like neck and neck cost wise to run something like this against a mini model and a best in class like flagship model. It's very interesting. Things are changing fast. Also of note is that O3 still has a 128k context window maximum. I thought they bumped it with O4 and I was wrong. Okay, they bumped it a little. They bumped it to 200k. But GBT 4.1 has 1 million, as do all of the Google models. All important things to think about as you reason about which of these models to use and what to use them for. More and more, it feels like you have to kind of be a power user to get the benefits of these things. Like once again, there's a reason that 4.0 is the default whenever you go to the chat GPT site. Most people should probably never change this. But us enthusiasts can benefit a lot from using these other tools the right way. These reasoning models will always have enough context. The problem is if you don't give them enough context, they will generate it themselves. So if you don't give it enough info to get to the right answer, it will just hallucinate a bunch of stuff and reason forever. That's why we saw this Yuchen post where it reasoned for four minutes when he said, hi, I'm Sam Altman. That's insane. That's hilarious. That's absurd. But that's because the model likes to parse through a bunch of information. That's what it's for. Which again, makes it weird that it has such a small context window. 200k token window input, 100k output, which is pretty absurdly large. But again, 200k in is small relative to a lot of these other models. So more than ever, there's this balance you have to strike of giving it the right information, giving it a lot of information, but also giving it the tools so it can access the additional information that it needs to be able to go get. But it's an interesting balance. And once again, to emphasize the fact that this model came out kind of late, May 31st, 2024, knowledge cut off. The knowledge the model has because of the data it was trained on is over a year old. Then again, so is O4Many. It seems like they're just training it all on the same knowledge. Interesting, actually. As Latent Space says, it's a really good model for using tools to do things and analyzing large amounts of data, but it's not so good at doing things directly itself. You ask it to just generate some code, probably not the best thing. You ask it what's wrong with a large amount of code, it'll probably do quite a bit better. They also call out that smart models like O3 Pro really need to be able to explore their environment using tools. This is an interesting point. O3 Pro feels very different from Opus and Gemini 2.5 Pro. Where Claude Opus feels big, but never showed me true clear signs of its bigness. O3 Pro's takes are just better. It feels like a completely different playing field. This is very interesting, and I'm excited to try it out and stuff like Cursor because, again, I have not seen much of a benefit from the bigger models there. I'm pretty sure I'm just rocking Sonnet 4 right now as my default. Yeah, Claude Force Sonnet is the model that I reach for at the moment, not even in max mode. I just use standard Claude Force Sonnet. It works pretty well for me. I was on 2.5 Pro, but I got more more annoyed about it not doing the thing it was supposed to. So I've used it less recently. I tried O3 a bit, and it was really smart, but it just took too long to do the things I wanted it to, and it was more expensive. It no longer is, so I'll give it another go. But if we look at what I actually use day-to-day, like in line for my command K, it's GPT 4.1 because I don't want a reasoning model to help me as I'm just quickly making a change. So if I highlight something and tell it, change how this looks or fix the layout shift or whatever, it's so nice having a non-reasoning model that just guns it out immediately. And 4.1 has been great for that. So that's where I'm at right now. But if O3 Pro is better at overhaul this portion of the code base, I could see it being really useful. I'm going to give it some honest shots tonight. And if I get enough interesting info, I'll leave a pinned comment telling you how I feel. It seems like once again, for these new reasoning models, system prompts should be minimized. Context should be maximized. Use cases should be long running, big tasks that they benefit from their intelligence for. Apparently, system prompts wildly shape behavior now more than ever. Leaves and bounds different from Anthropoc and Gemini, where cloud opus feels big, but doesn't really show it. Yep. And it seems like OpenAI is really going down this vertical reinforced learning path, stuff like deep research and codecs, trying to make it so they can do smarter things with the tools. And I totally agree. In a lot of ways, OpenAI was kind of late to this tool calling world. But OpenAI is rarely late to things. And now it feels like they are leaping ahead. It really does. I think we'll be in a position very soon where Claude's advantage on tool calls is fully closed. And the gap is now just how we feel about them as developers. And more than ever, it does feel like OpenAI is taking their competition seriously and going out of their way to never be number two. With Gemini, it got close for a brief moment. And honestly, if the Google stuff was more stable, I would have said they were at the number one position. It was very, very close. But with O3 getting cheaper, O3 Pro finally shipping, and O4 Mini being as insanely cheap and effective as it is, and all of these being on APIs that are consistent, having tool call protocols that work reliably, the embrace of MCP, and all these other things, I still feel like OpenAI models are the ones that I can just use and get what I'm looking for. Are they as good at coding in my editor? Maybe. I haven't used them as much recently, but I'll let you know in a bit. But overall, OpenAI has managed to maintain their lead. And once again, these ships put them in first place. I'm curious how you guys feel. Is this AI race getting less interesting, or are these leaps still exciting to you? I'm still hyped, especially when the price gets cheaper so I can give you guys better stuff on T3 chat, but I want to know how y'all feel. Let me know in the comments. Until next time, keep prompting.