← Back to video archive

Airdroplet AI summary

Mistral Reasoning Model, Gemini 2.5 Update, FLUX.1 Kontext [Max], Meta's Spending Spree

June 12, 2025Matthew BermanAI score 10060,514 views

Watch original on YouTube ↗

AI-generated summary

This week brought a flurry of exciting AI news, covering breakthroughs in reasoning models, hyper-realistic voice AI, advancements in text-to-video, and significant shifts in the competitive landscape of major tech players. We saw a new contender for the fastest AI model, surprisingly human-like voice capabilities, and Meta making a bold move to secure top AI talent and data.

Here's a breakdown of the key developments:

  • Mistral's Magistral Reasoning Model is a Game Changer: Mistral dropped their first reasoning model, Magistral, and it's incredibly fast – "by far the fastest reasoning model" available, easily outperforming even Gemini 2.5 Pro. There are two versions: Magistral Small, a 24 billion parameter open-source model that you can download and run on most consumer-grade computers right now, and Magistral Medium, a more powerful enterprise version. Magistral Medium scored an impressive 73.6% on AME 2024, with Magistral Small close behind at 70%. It handles chain-of-thought reasoning across multiple languages and runs at 10 times the speed of competitors, thinking for 5.3 seconds compared to OpenAI's 17 seconds in one comparison. You can try it for free on Mistral Lit Chat app.

  • Eleven Labs v3 Alpha Delivers Emotional Text-to-Speech: Eleven Labs released the V3 Alpha of their text-to-speech model, which is described as their "most expressive, most emotional voice model to date." It boasts incredible clarity and can now produce subtle vocalizations like whispers and various emotional tones. While the "laugh upgrade" was a bit unsettling, the overall realism is striking. You also get much more control over voice nuances by adding tags like "excitedly," "surprised," or "cautiously."

  • OpenAI's Voice Model is Almost Too Human: OpenAI also launched an upgraded voice mode that is "scary realistic." It includes natural human speech patterns, complete with "ums" and realistic pauses, making it incredibly lifelike. While impressive, there's a personal preference for it to sound "a little more AI-like," as the human imperfections can be a bit much. It's so good that it's become a go-to for learning while driving.

  • Gemini 2.5 Pro Gets Another Boost: Less than a week after its previous update, Gemini 2.5 Pro received another significant upgrade, making it "the best Gemini 2.5 Pro model yet." This version shows even better performance on benchmarks, with a 24-point ELO jump in LM arena and a 35-point jump on web dev arena, maintaining its number one spot. It continues to excel in coding, leading on difficult benchmarks like Ader Polyglot, and remains the top choice for complex coding challenges. It's free to use via AI Studio by Google.

  • Google's Veo3 Text-to-Video Now Faster and Cheaper: Google's popular text-to-video AI model, Veo3, now has a new "fast" version. This option is significantly quicker and is only one-fifth the price of the original Veo3, making it more accessible for experimenting with AI-generated videos.

  • Meta's Massive AI Play: Acquiring Scale AI and Chasing Talent: This week's biggest news involved Meta's strategic shift in the AI race. Meta made a substantial $14 billion investment to acquire a 49% stake in Scale AI, essentially securing control without undergoing the full regulatory hurdles of an outright acquisition. Scale AI is renowned for building powerful data labeling and annotation engines, providing Meta with access to high-quality, rich data essential for AI development. Adding to this, Meta hired Scale AI CEO Alex Wang to lead a new "super intelligence team" directly overseen by Mark Zuckerberg. Zuckerberg is reportedly hand-picking 50 of the "top AI minds" in the industry to build super intelligence, indicating a desire to accelerate Meta's AI progress and potentially a frustration with current internal efforts. The competition for this talent is fierce, with Meta reportedly offering "insane" compensation packages, including over $10 million per year in cash.

  • Dia Browser Aims to be AI-Native: The creators of the ARC browser introduced Dia, an "AI-native browser," getting a jump on Perplexity's upcoming Comet browser. Dia's key selling point is the ability to "chat with your tabs," using AI to interact across multiple open web pages. However, it's questioned whether features like inline copy editing, grammar checks, or summarization are truly innovative for a browser, as many of these functionalities are already built into native tools like Gmail, Google Docs, and Notion. The hope is that having everything "all in one place" might prove beneficial. You can join the waitlist to try it.

  • FLUX.1 Kontext [Max] for Top-Tier Text-to-Image Generation: The FLUX.1 Kontext Max model is being hailed as "one of the best text-to-image models on the planet," rivaling Google's Imagine 4. Developed by Black Forest Labs, its image generation quality is impressive, scoring very close to top models like GPT-4.0 and Seedream Recraft V3 in benchmarks. While the Max and Pro versions are only available via API, Black Forest Labs is developing FLUX.1 Kontext Dev, a 12 billion parameter diffusion model, which they plan to open-source soon. Image examples show highly detailed and impressive outputs, though minor imperfections can still be found.

Video transcript

Open transcript
So much news happened this past week. Let's go over it all. First, from Mistral, they released their first reasoning model and they open sourced the smaller version of it. And here's the thing, it is by far the fastest reasoning model I have ever used. I thought Gemini 2.5 Pro was fast. This leaves it in the dust. So here's what you need to know. We're releasing the model in two variants, Magistral Small, a 24 billion parameter open source version of Magistral Medium, a more powerful enterprise version. This is something you can download and run on your computer right now. 24 billion parameters is relatively small. And when it gets quantized down to even smaller sizes, you'll be able to run it on most consumer grade computers. Magistral Medium scored a 73.6% on AME 2024 and 90% with majority voting at 64, meaning 64 attempts. Magistral Small scored 70%, so nearly the Magistral Medium and 83% respectively. Magistral's chain of thought works across global languages and alphabets, and it runs at 10 times the speed compared to most competitors. Just to show you how fast it is, on the left is Magistral, on the right is one of OpenAI's models, although it's not exactly clear which one it is. Watch how fast it is though. Look at that. Look at how fast it's outputting, plus we get more details in the chain of thought. So thought for 5.3 seconds, and the OpenAI model thought for 17 seconds, and is still outputting the final answer. So give it a try, download it, or use it on the Mistral Lit Chat app. It is free, so let me know what you think. Next, Eleven Labs released V3 Alpha of their text-to-speech model. This is the most expressive, most emotional voice model to date for Eleven Labs. Listen. Hey Jessica, have you tried the new Eleven V3? I just got it. The clarity is amazing. I can actually do whispers now, like this. Ooh, fancy. Check this out. I can do full Shakespeare now. To be or not to be, that is the question. Nice. Though I'm more excited about the laugh upgrade. Listen to this. Okay, besides for that incredibly creepy laugh, everything else sounded really good. And these voice models are getting incredibly realistic, almost too human-like. In fact, OpenAI just released an upgrade to their voice mode that sounds really just too much like a human. There are too many ums and stutters and mistakes, and I would actually prefer it to sound a little bit more AI-like. But I'll talk about that later in this video. And so you now have a lot more control over what the voices and exchanges sound like. So excitedly, jumping in, meaning interrupting, surprised, overlapping, cautiously. So you can add these little tags, which is really cool. So check it out. New V3 Alpha from Eleven Labs. And again, speaking of voice, OpenAI released their new voice mode, and I have been testing it out. And it is really good, like scary realistic. Let me give you an example of what that sounds like. Tell me about the semiconductor industry. The semiconductor industry is basically the backbone of modern electronics. It involves the design and manufacturing of semiconductor devices, which are essential components in everything from smartphones to cars to industrial machinery. The industry is super dynamic with constant innovation and competition. And it's really global with key players in the US, South Korea, Taiwan and Europe. All right, that sounds incredibly good. I don't know if you heard it, but there were a couple ums. There were specific pauses that sounded realistic. Even when she was listing off lists of things, the list sounds, I mean, I'm thinking about how I'm talking right now because that sounded so much like how I talk. I'm glad you thought it sounded realistic. It's always interesting to hear that feedback. If you ever want me to adjust how I speak or if there's a particular style you prefer. Okay. Okay. Be quiet. All right. So she was still listening, but yeah, very, very good. And I've gotten into the habit when I'm driving around to just start talking to her and have her teach me things. Let me know if you do that. All right. Next, this came out less than a week ago, but it already seems like old news. Gemini 2.5 Pro has a brand new version. This version ekes out even more on different benchmarks. So definitely the best Gemini 2.5 Pro model yet. So it has a 24 point ELO jump in LM arena, maintaining the number one spot at 1470, a 35 point ELO jump to lead on web dev arena at 1443. It continues to excel at coding, leading on difficult coding benchmarks like Ader Polyglot. So still to this day, Gemini 2.5 Pro is my favorite coding model, at least when I'm going directly to it and asking it to solve things like the Rubik's cube test. So check out the new model. It's free to use AI studio by Google. Next, another quick Google update VO, the incredibly popular text to video AI model from Google has a new fast version. This new fast option is one fifth of the price of VO3 and is significantly faster as well. Hence the name. I love playing around with the VO videos. So I'm definitely going to be trying this out. And thanks to the sponsor of this video, OutSkill. OutSkill is a live two day AI training program for professionals, founders, and executives. Through this live two day program, you will master the skills of AI, including the basics of generative AI, automations, building AI agents, image and video generations, generating full fledged websites, and more. The two day training happens Saturdays and Sundays from 11am to 7pm Eastern. And there's an initial kickoff Friday 10am. Two days, 16 hours, five sessions. You will learn so much. 50,000 professionals have already attended these sessions over the last six months, and they've landed consulting gigs, built AI products, or just upskilled themselves in their existing roles. They also offer live Q&A sessions with mentors so you can ask the questions and clear up any doubt that you might have and clarify any topics that might be confusing. So check out OutSkill. I'll drop a link down below. It's free for the first 1000 people who register. And thanks again to OutSkill. Now back to the video. All right. And the big news this week, Meta has made a major investment in scale AI and is shaking up their AI team. So Meta forming new AI lab helmed by scale AI CEO Alex Wang report says, and yes, the report seemed to be accurate. Zuckerberg feeling like Meta is falling behind in the AI race made a $14 billion investment in scale AI for 49% of the company and hired the CEO. So the CEO is no longer the CEO of scale AI. He is now leading up the newly formed super intelligence team that apparently is being handpicked by Zuck himself. Zuck is looking for 50 of the top AI minds in the industry to build super intelligence. It seems like maybe Jan LeCun is not delivering to what Zuck's expectations are. And if taking a 49% stake sounds weird, like why didn't they just acquire the whole company? Well, they probably didn't want to go through the regulatory hurdles to actually do that. So this roundabout way of acquiring a minority, but the majority of the minority stake 49% in a company is kind of the way around that. Google has done that. Microsoft did that with open AI. So this seems to be the trend in acquiring companies. And if you're not familiar with scale AI, they essentially built an entire engine for data labeling and annotation for AI companies. Really powerful stuff, really good, high quality, rich data. And now Meta gets all of it. And yeah, Zuck is going hard after the top minds in the AI industry. This is according to Didi. This is not verified whatsoever, but it's true. The meta offers for the super intelligence team are actually insane. Zuck is personally negotiating 10 million plus dollars per year in cold, hard, liquid money. I've never seen anything like it. So every major AI company is competing for the same finite set of talent. And it is a complete cutthroat race. Next, from the company that makes the ARC browser, they now have the Dia browser, which is an AI native browser. Getting ahead of Perplexity, who is launching their own browser, Comet, very soon. This browser is emphasizing the fact that you can quote unquote chat with your tabs. Basically, you have a bunch of tabs open and you can use AI to chat across them. I personally don't know what's so special about that, but I haven't tried it. I'm going to give them the benefit of the doubt and I want to test it out. I will let you know. So here's an example, an inline copy editor. So highlight a part of your Gmail email, say, make this sound more confident and boom. Gmail already does this. So I don't see what's so special about that. Here, make sure I don't sound stupid. Any typos or grammar issues. Again, this is all built in natively into Google Docs. Here's what looks to be Notion. Summarize for Slack. Okay. Summarization. Again, already done with Notion. So all of these things are already done with the native tools, but maybe it's nice just to have it all in one place. I don't know yet. So if you want to give it a try, go ahead, join the wait list. Next, Artificial Analysis says the Flux 1 Context Max model is one of the best text-to-image models on the planet. And not only that, it is open source. So not only an impressive image editing model, it's also one of the best text-to-image models, rivaling Google's Imagine 4 in the Artificial Analysis image arena. This is by Black Forest Labs, released just about a week ago. Now, the Max and Pro versions are not open wait, so keep that in mind. These are only available via the API or other partner providers. Black Forest Labs are also developing Flux 1 Context Dev, a 12 billion parameter diffusion image editing model that they plan to make open waits soon. It's currently in private beta release. So OpenAI GPT-4.0, still taking the top spot. Then Seedream Recraft V3, Imagine 4 Ultra and Preview, then Flux 1 Context Max. So very close, very good model. Here are some example images from this new model. Hovering Antarctic Research Base. Here's Flux 1 Context Max, Flux 1.1 Pro Ultra. Here is GPT-4.0 and Seedream 3.0. And they're all really good. This one looks more like an illustration, but yeah, they're all really good. Here's another example. Neon Lit Alley in Tokyo bustling with animated crowds under a rainy sky in anime style. This is Flux 1 Context Max, Flux 1.1 Pro Ultra, GPT-4.0 and Seedream. Now, again, all four of them look really good. I'd say this one probably is my favorite because it has the most detail, Flux 1.1 Pro. But all of them, again, really, really good. Here's another one. Young Cartoon Pirate Adventurer setting sail on the high seas. So this one by Flux 1 Context Max, very good. Although the patch over the eye is kind of broken. Here's 1.1 Pro Ultra, which is very good. The only mistake I see here is the water kind of looks like it's coming out of the boat. Here's GPT-4.0. The pirate's leg is kind of overlapping the boat. And Seedream 3.0. I don't see any mistakes with this one. So that's all the news for today. If you enjoyed this video, please consider giving a like and subscribe. And I'll see you in the next one.