← Back to video archive

Airdroplet AI summary

AI NEWS: GPT-4o Major Updates, Gemini 2.5 Pro, New DeepSeek, MCP Everywhere, New Image Models

March 29, 2025Matthew BermanAI score 10095,081 views

Watch original on YouTube ↗

AI-generated summary

Here's a rundown of the latest AI news, touching on significant updates to major models, new benchmarks, and interesting industry trends. We cover how older models are getting surprisingly powerful upgrades, new contenders are shaking up the coding and image generation scenes, and a crucial protocol for connecting AI agents to tools is becoming the standard everywhere. Plus, there's a peek into the booming financials of leading AI companies and some cool open-source developments from overseas.

  • GPT-4o Got Some Serious Muscle: GPT-4o received major updates, making it surprisingly powerful again. It's now considered the best model for image generation and, according to independent benchmarks (like Artificial Analysis), the leading "non-reasoning" model for coding, surpassing strong competitors like Claude 3.7 Sonnet and Gemini 2.0 Flash.
  • Why Update an Older Model? The reason OpenAI is pushing updates to 4o seems linked to a concept called Javon's paradox – as things get cheaper, demand increases. Even OpenAI, a giant with massive funding, is facing GPU shortages needed for training newer, larger models like 4.5. So, they're optimizing and enhancing the existing 4o, which is likely more cost-effective to run at scale.
  • More 4o Improvements: Beyond coding and image generation, the updated GPT-4o is better at following complex instructions with multiple requests, handling technical coding problems, and shows improved intuition and creativity. On the lighter side, it uses fewer emojis now.
  • Availability & Issues: The updated GPT-4o is currently available to all paid users, with free users getting access over the next few weeks. However, its image generation capability is already seeing rate limits due to unexpectedly high demand, and the overall speed for normal queries feels "unusably slow" right now, which is a problem because speed is a crucial factor for a preferred model.
  • Gemini 2.5 Pro is a Coding Beast: This was another huge news item. Gemini 2.5 Pro is incredibly good at coding and, importantly, it's "super fast." It's called a "full thinking model" and is considered the "best coding model ever used" by the speaker, who finds speed especially vital for agentic and coding tasks.
  • Massive Context Window: A key feature of Gemini 2.5 Pro is its massive one million token context window, about 10 times larger than Claude 3.7's. This is exciting because a larger context window should allow the model to understand entire codebases better, which is being actively tested.
  • Gemini 2.5 Pro Availability: Good news for developers using specific tools: Gemini 2.5 Pro is now available in Windsurf and Cursor, making it easier to test and integrate into coding workflows.
  • New DeepSeek V3 Checkpoint: It was a big week for models! A new version (checkpoint) of DeepSeek V3 was released quietly. While not a completely new model, this update significantly improves its performance, especially excelling at coding, math, and logic tasks.
  • DeepSeek V3 is Fast & Open: DeepSeek V3 is also fast and, importantly, it's open-source and uses a very permissive MIT license. This is great because it allows anyone to download and potentially run it (though it's a large model, which might be a challenge locally) or use it via inference providers.
  • DeepSeek V3 Benchmarks: Benchmarks show the new DeepSeek V3 performing extremely well against frontier models like the previous DeepSeek V3, QwenMax, GPT 4.5, and Claude Sonnet 3.7, especially dominating in math tests like AIME 2024, despite many of those competitors being closed-source.
  • ARC-AGI2 Benchmark Released: The ARC Prize organization launched ARC-AGI 2, their new benchmark designed specifically to test models' "AGI-ness." These tasks require abstract reasoning and the ability to apply understanding from one context to another, something humans find relatively easy but AI models struggle with.
  • AI vs. Human Performance on AGI Benchmarks: The current scores highlight the gap: The best AI models (like O3 Low) score very low on ARC-AGI 2 (only 4%), whereas humans achieve a perfect 100%. This significant difference is seen as proof that the benchmark is effectively testing true generalization capabilities beyond current AI strengths. There's still a million-dollar prize for solving the ARC AGI tasks.
  • MCP is Becoming the Standard: The Model Context Protocol (MCP), a way for AI agents to easily connect and use various tools (like Zapier actions), is rapidly gaining traction and seems to be becoming the industry standard.
  • Widespread MCP Adoption: Zapier (which offers connections to 10,000+ tools) announced its adoption of MCP, allowing agents to directly access its vast library of actions. OpenAI also adopted MCP as part of its agents API, enabling agents to use tools via the protocol. Microsoft is integrating MCP into Copilot Studio. This widespread adoption means wherever you run your agents, you'll likely be able to use MCP to give them tool access.
  • Anthropic's Influence: The speaker notes that while MCP is an industry standard, Anthropic gets credit for setting it, giving them some influence in its development.
  • Text-to-Image is Hot Right Now: This week saw significant advancements in text-to-image generation beyond GPT-4o's improvements.
  • Reve Image 1.0: Reeve AI launched its text-to-image model, Reve Image 1.0, which looks really good based on examples shown, featuring accurate text rendering and diverse styles. It ranks highly in quality based on user votes in artificial analysis rankings.
  • Ideogram 3.0: Ideogram also released 3.0, which looks phenomenal. While Ideogram claims the highest ELO rating for quality, the key takeaway is the high degree of control offered through features like remixing, upscaling, and style preferences, allowing users to create beautiful, hyper-realistic images with lots of customization.
  • OpenAI is Making Bank: Financially, OpenAI is booming. Sources report they expect revenue to triple to $12.7 billion this year, although they are still currently losing money overall.
  • AI is Not a Fad: This massive revenue growth is seen as strong evidence that AI is definitely not a fad and that the significant investment in the field is resulting in substantial value. The speaker feels this personally, using AI tools extensively, and believes the main barrier to wider adoption is people not knowing what's possible or how to use the tools effectively – a problem the speaker is personally trying to solve through education.
  • OpenAI Leadership Changes: Some recent C-suite changes were noted: Sam Altman is shifting focus away from daily operations to concentrate more on research and product, while operating chief Brad Lightcap will take on a larger role overseeing business and day-to-day activities.
  • Massive Valuation: SoftBank is reportedly set to invest $40 billion in OpenAI, valuing the company at an astounding $260 billion, which would make it one of the most valuable private companies ever.
  • Quen Qvq Max - Visual Reasoning Powerhouse: Chinese company Quen released Qvq Max, an open-source visual reasoning model that can understand and reason with information from images and videos. This open-sourcing trend from Chinese companies is seen as a positive development, allowing users to customize and run the AI themselves.
  • Quen Qvq Max Capabilities: This model can handle complex tasks like analyzing relationships between scenes in multiple images, solving math problems, and generating code or artistic creations based on visual input. It is a "thinking model with vision capability."
  • Availability Challenge: The main drawback for US users is that Quen's services often require a Chinese phone number. However, there's hope that this powerful model will soon be available via third-party inference providers that US users can access, or potentially through quantized versions that are easier to run locally, as the full model is quite large.

Video transcript

Open transcript
You want to hear something crazy? GPT-40 got some massive updates. In fact, it is not only the best at image generation now, but it's also the number one non-reasoning coder on the market. And this is verified by independent benchmarks. Look at this. Artificial analysis. Today's GPT-40 update is actually big. It leapfrogs CLAWD 3.7 Sonnet non-reasoning and Gemini 2.0 Flash in our intelligence index and is now the leading non-reasoning model for coding. Here's the intelligence index. This is not the coding index. This is the intelligence. It goes from 41 back in November, 2024. That version of the model leapfrogs all the way up to 50 on their score, just behind DeepSeq V3, the recent model that just came out. It's kind of nuts. Why are they putting so much time and effort into an older model like 4.0? Well, it turns out there's actually a very explicit reason. Anybody who thought that the DeepSeq R1 phenomenon meant that everybody was going to start spending less on compute. Boy, you really should have listened to the experts with Javon's paradox. As things get cheaper, we want more of it. And that's what we're seeing. They can't even get enough GPUs to fine tune and improve 4.5. We're not talking about some small startup. This is OpenAI, partnered with Microsoft, a multi-billion dollar company that has raised billions and billions of dollars. They can't even get enough of these chips. But those aren't all the improvements. GPT-4-0 got better at following detailed instructions, especially prompts containing multiple requests, improved capability to tackle complex technical encoding problems, improved intuition and creativity, fewer emojis. The updated GPT-4-0 is available now to all paid users. Free users will see it over the next few weeks. Now, here's the thing. They're already having to put in place rate limits on the image generation capability of 4.0 because it far exceeded even their expectations. And I'm already seeing it's super slow. I started noticing this a few weeks ago. GPT-4-0 for normal queries is almost unusably slow at this point. And they really need to speed that up because one of the factors that I look for in my go-to model is speed. And next, the big news this week was also Gemini 2.5 Pro coming out of the gates swinging. It is incredible at coding. This is a full thinking model and incredibly fast. I made a full video about it already, so I'm not going to dive too deep into it, but it is the best coding model I have ever used. And again, I want to reemphasize super fast. And I think people under appreciate how important speed is, especially with agentic use cases and just as importantly, coding use cases. And now we have the best new coding model in the world and vibe coding is really all I'm doing in my free time right now. So of course I was excited to see our next story. Gemini 2.5 is now available in windsurf. And not only that, it's also available in cursor. So I'm going to be testing it thoroughly. The great thing about 2.5 Pro is that it has a million token context window, which is about 10 times what cloud 3.7 has. So I really want to see how that affects its ability to understand the code base as a whole, because I certainly have less than a million tokens of code on my projects that I'm working on right now. So I'm going to be testing it out. I will report back. Next, this really was an incredible week for new models. There is a new version of DeepSeq v3. It came out just a few days ago. They really made no splash about the announcement. It is a new checkpoint in the v3 series. So it's not a completely new model, but it really excels at coding and math and logic, and it is fast and open source and you can download it, although it is a massive model. So you may have trouble running it locally. So here are the scores. DeepSeq v3 and the Striped Dark Blue versus the previous DeepSeq v3, QuenMax GPT 4.5 and Cloud Sonnet 3.7. Now, it would be interesting to see this versus 4.0 new, but look at that. Again, DeepSeq v3, a non-thinking model compared to all of these other Frontier models and GPT 4.5 and Cloud 3.7 are both closed source and it performs extremely well, especially at math. Look at that AIME 2024 score absolutely dominating the rest of the models in this comparison. And it is open weights and they also switched to an MIT license, which is very permissive. So go download it and have fun. And the ARK Prize company has now released ARK AGI 2. This is their new benchmark to test the AGI-ness of models. And so we can see the scores right here. O3 Low currently has the highest score. And on ARK AGI 1, it has a 75.7% score. For ARK AGI 2, only 4%. And look at this actually, $200 per task for O3 Low. And look at the number one, the human panel. For ARK AGI 1, they got 98%, so not even perfect. And then ARK AGI 2, they got 100%. And I love this. The fact that humans can score a perfect score on this test, but the best of the best models score so low, that to me is the perfect benchmark for testing AGI. And the cost per task for humans, $17. Now look down here. O1 high, $4.45 and down from there. And if you're not familiar with the ARK Prize benchmarks, they are essentially tasks that require you to take an understanding of one thing and extrapolate it out to understand other things. So here's an example. We have an input here, and all we do is just look at it. We have a 30 by 30 grid. We obviously have kind of this delineator right here. We have a bunch of colors and squares here. And then we have a bunch of kind of gray blobs in the middle. And so the point is you're trying to look for patterns between the two example inputs and outputs, and then recreate it over here. And so if you want to try it out, you can. I'll drop a link down below. And so the point is, this is supposed to be easy for humans, but really hard for AI. And once again, there's a million dollar prize. So congratulations to ARK Prize for the new benchmark. Next, Zapier announced their own MCP. That's basically like getting 10,000 tools all at the same time for MCP. All you got to do is sign up, configure the apps that you want, and then they'll give you the MCP server URL. I'm personally a big fan of Zapier. I've been using them for years and years. It's great for automation. And the fact that now I can connect my agents and my AI directly to it, just fantastic. But that's not all. OpenAI just adopted MCP as well. It's quickly becoming clear that MCP is the standard. And now part of their agents API, you can now use MCP to give your agents tools. And that's not all. MCP has now been adopted by Microsoft. Introducing model context protocol in Copilot Studio. So it seems wherever your agents live, that's where you're going to be able to use MCP. And this is a big win for Anthropic who set the standard. Even though this is an industry-wide standard, whoever creates the standard definitely gets to put their thumb on the scale. And this was definitely the week of text to image generation. Not only did we have 4.0 absolutely dominating the headlines with text to image, but Reeve AI came out with their text to image and it looks really good. Look at all of these. The text is all accurate. All of the different styles is great. Look at this realistic piece of steak on a salt block nature, more artistic. And according to the artificial analysis rankings, Reeve image 1.0 ranks highly in quality based on a hundred thousand user votes and half moon is Reeve. And like I said, this was the week of text to image. Ideogram launched 3.0 and it is also looking phenomenal. Now, of course, Ideogram says they are scoring the highest in the ELO rating, but you could do things like remix upscale style preference and so on. So highly controllable, which is great. And just look at these images. They're beautiful. They're hyper realistic and you get lots of control over them. So although 4.0 seemingly got a lot of the press love this week, there are multiple other fantastic image models that just got released this week. So check them out. And now back to OpenAI. Turns out they're printing money. Now, although they're still losing money, they are making a ton. According to CNBC, OpenAI expects revenue will triple to 12.7 billion this year, sources say. AI is not slowing down. Occasionally, I get the question, do you think this is a fad? Do you think all of this spend is not going to result in value? Not a chance. I am obviously extremely biased. I'm dedicating my life to this. But for the amount of value, just personally, a data point of one that I get out of all of these AI tools is tremendous. And my team uses these tools every day. My family uses these tools every day, all day. They've been exposed to them because I've told them about them. I've told them how to use them. I actually think the problem is an education problem. People just don't know what's possible. And that's the problem I'm trying to solve. I want everybody in the world to understand how to use these AI tools and how to get the most out of them. That is my mission. So a few other notes about OpenAI. So earlier this week, OpenAI announced some key changes in the C-suite. Sam Altman will shift his focus away from day-to-day operations and focus more on research and product, which is very interesting. While operating chief Brad Lightcap's role will expand to oversee business and day-to-day operations. SoftBank just last month was set to invest $40 billion in OpenAI at a $260 billion valuation, making it one of the most valuable private companies ever. Next, Quen released Qvq Max. Think with evidence. This is a visual reasoning model. And it's open source. These Chinese companies open sourcing everything are really making these US companies look bad. And you know what? I'm all for it. I really appreciate open sourcing it because once you can get it yourself, you can make the AI any way you want. This model can not only understand the content and images and videos, but also analyze and reason with this information to provide solutions from math problems to everyday questions from programming code to artistic creation. Qvq Max has demonstrated impressive capabilities. So here's an example of what it can do. Let's upload these two images and the scenes depicted in these two pictures. What is the relationship? So it is a thinking model with vision capability. So you can see it's thinking right now. And there you go. It gives you an answer. Now, the only problem is Quen is not open to US users. You have to have a Chinese-based phone number to use it, but that's okay. We're going to get this model on inference providers that allow for US users. And you can also download it yourself. Now, this is also a relatively large model. So yeah, you're probably going to struggle running it. Maybe we're going to get quantized versions of it, but we want to run the full thing and hopefully some of the awesome providers will do it soon. So that's all the news for this week. If you enjoyed this video, please consider giving a like and subscribe and I'll see you in the next one.