← Back to video archive

Airdroplet AI summary

Mathematicians STUNNED as o4-mini answers the world's hardest math problems...

June 9, 2025Wes RothAI score 10048,846 views

Watch original on YouTube ↗

AI-generated summary

Alright, let's dive into what's happening in the world of AI and math, because it's pretty wild! This video unpacks the stunning capabilities of new AI models, particularly OpenAI's O4 Mini, in solving incredibly difficult math problems, and how this is pushing the boundaries of what we thought AI could do. It really makes you think about how our relationship with AI is evolving from simple tools to genuine collaborators, and possibly even independent discoverers.

Here's the lowdown:

  • The "Secret Math Meeting" – AI vs. Human Brains: Imagine 30 of the world's best mathematicians gathering secretly to try and stump an AI. They were absolutely stunned to find O4 Mini could solve some of the hardest solvable problems. It wasn't just solving them; it was showing its reasoning in real-time, even tackling open questions in number theory that would be PhD-level for a human. For instance, one mathematician, Ken Ono, watched in silent amazement as O4 Mini unfurled a solution in minutes, first researching related literature, then trying a simpler "toy version" of the problem, and finally presenting a correct (and apparently "sassy") solution.
  • O4 Mini's "Sass" and Confidence: The AI wasn't just smart; it had attitude! Its solutions came with confident, almost cheeky remarks, like "no citation necessary because the mystery number was computed by me." This led to a concern about how much trust humans might place in AI, especially since these models, even when wrong, can present information with an air of complete confidence. It means we have to be careful and always verify.
  • Tier 4 Problems and Beyond: The problems O4 Mini was tackling were considered "tier four," meaning they're at the very top of human ability. The discussion at the meeting quickly shifted to "tier five" problems – those even the best mathematicians can't solve. The question now is, what happens when AI becomes superhuman at those? It's clear that the roles of mathematicians are about to change dramatically, possibly shifting from direct problem-solving to overseeing and collaborating with AI.
  • The Frontier Math Benchmark & OpenAI's Involvement: To measure AI's progress in math, a new benchmark called "Frontier Math" was created because AI was already acing existing human-level tests. OpenAI actually commissioned Epic AI, a nonprofit, to create 300 unpublished math questions for this benchmark. This raised some eyebrows, as it means OpenAI had access to much of the data, though 50 questions were kept as a "holdout set" that the models had never seen. This ensures legitimate testing, proving AI can solve truly novel problems.
  • Google DeepMind's AlphaProof/AlphaGeometry - Almost Gold: It's not just OpenAI; Google DeepMind's AlphaProof and AlphaGeometry also showcased incredible math skills, achieving a silver medal in the International Mathematical Olympiad (IMO) – just one point shy of gold! These tests are rigorously fair, with problems kept secret until the competition day.
  • The "Flawed Reasoning" Problem: A mathematician named Jasper, who was at the secret meeting, clarified that they used O4 Mini in its "high thinking mode." While it solved most problems, he noted that sometimes the AI would arrive at the correct numerical answer despite its reasoning being "occasionally incorrect." This is a known challenge in AI training where models are rewarded for correct outputs but the internal logic isn't always fully verifiable or sound. It's a tough problem to fix if you're only reinforcing the final answer and not the intricate steps of reasoning.
  • AI's Strengths and Weaknesses in Math (Currently): While AI excels at gathering relevant literature and drafting initial solutions, it still struggles with deep reasoning, especially when it needs to synthesize complex ideas from different sources into a novel computational method. So, it's not yet generating completely new mathematical theories on its own. Human oversight remains crucial for verification and for pushing the boundaries of true synthesis.
  • The Future: AI as Collaborator and Independent Discoverer: The prediction is that in the next year or two, AI will transition from assisting mathematicians to collaborating with them in discovering new theories and solving open problems. Eventually, it could even work independently to push the frontiers of mathematics and other scientific fields.
  • Recursive Self-Improvement: Alpha Evolve & Darwin Gödel Machine: This is where things get really exciting. Google's Alpha Evolve, powered by Gemini, is already discovering advanced algorithms that optimize Google's own data centers, saving significant resources. It can even improve its own training, hinting at recursive self-improvement. The Darwin Gödel Machine takes this further, acting as a self-improving coding agent that uses an evolutionary search process: it generates many potential solutions, tests them, and builds upon the promising ones, eventually outperforming human-coded agents.
  • The Power of Iteration and Feedback: The presenter thinks the math symposium's results, while impressive, might have been even more "staggering" if they had incorporated the iterative, feedback-driven approach of Alpha Evolve or the Darwin Gödel Machine. These systems can generate thousands of outputs and continuously refine solutions, whereas the symposium likely only tested one output at a time. Automating verification and synthesis steps could supercharge these AI systems, making them incredibly powerful.
  • The "Religion of Justism": The video closes with a powerful point from Scott Aaronson, highlighting the common human tendency to constantly deflate AI's achievements by saying it "just" does X (e.g., "it's just a stochastic parrot," "it's just a next token predictor"). He challenges this by asking, "What are you just a?" reminding us that we, too, could be reduced to a "bundle of neurons and synapses." This "justism" ignores the real-world impact and capabilities of AI, which are already changing civilization.

Video transcript

Open transcript
It was, okay, they gave this giant litany. Look, GPT does not interpret sentences. It seems to interpret them. It does not learn. It seems to learn. It does not judge moral questions. It seems to judge moral questions. And so I just responded to this. I said, that's great. And it won't change civilization. It will seem to change it. So this paper got published in the last couple of days. At a secret math meeting, researchers struggled to outsmart AI. The researchers were stunned to discover it was capable of answering some of the world's hardest solvable problems. Now, of course, not everyone agrees with that article. There's also a post from one of the people there, one of the best mathematicians in the world, as he was described, Jasper. So he was there at that math symposium, and he's going to say what actually happened. As he says, some parts were a bit exaggerated. So this by no means reverses the whole thing. There was just some bombastic language that was used, if you will. So we'll come back to that. Now, the idea that AI is getting really good at math shouldn't come as a surprise. We have the frontier math, a brand new benchmark that basically had to be created because the AI is getting so good at math that sort of the normal human level problems, those benchmarks are being saturated, meaning that they're approaching 100%. So we needed to create something that's, you know, on the frontier, frontier math. You know, here's an example of a question on there, right? A recursive construction on large permutations. I mean, there it is. You're welcome to read it. For our audio listeners, it's, for most people, is just absolute gibberish. The first sentence goes, let W be the set of finite words with all distinct letters over the alphabet of positive integers. And it gets a lot more complicated after that. But these are sort of the benchmarks that we're using to test the next generation AI models. Now, as you'll see, OpenAI does have some connection to frontier math. Either they funded it or they worked on it somehow. We'll get to that just in a little bit. So for some people, that was a bit of a red flag. But we also have tons of other examples of AI models being very good at math. Google DeepMind, they had their alpha proof and alpha geometry. It achieved a silver medal on the IMO, International Mathematical Olympiad. And that's underselling it a little bit. It was one point away from gold. And there, of course, all the problems are kept in the strictest secrecy until the day of the competition. So here, it was actually the people that were running the IMO that helped facilitate this test. So everything was done, you know, legitimately. This is a legitimate result. Other benchmarks like the AIME, you know, these models get tested on those for the brand new year before those answers are available publicly. So again, there are tests where we see these models do extremely well in situations where we're very confident that they haven't seen those specific problems before. So this article begins. On a weekend in mid-May, a clandestine mathematical conclave convened, which is just a great, great sentence. But basically, 30 of the world's most renowned mathematicians traveled to this location and faced off in a showdown with a reasoning chatbot. The researchers were stunned to discover it was capable of answering some of the world's hardest solvable problems. I have colleagues who literally said these models are approaching mathematical genius, says Ken Ono, a mathematician at the University of Virginia and a leader and judge at the meeting. And the chatbot in question is powered by the O4 Mini. And this is where we get to Epic AI and the Frontier Math benchmark. So to track the progress of O4 Mini, OpenAI previously tasked Epic AI, a nonprofit that benchmarks large language models, to come up with 300 math questions whose solutions have not yet been published. So in the past, when these models were asked these very hard questions, very, very hard questions, they were able to answer less than 2% of them, showing that these LLMs lack the ability to reason. But O4 Mini would prove to be very different. So importantly, this is from Epic AI, right? So they're kind of clarifying what OpenAI has access to, what they don't have access to. So they're saying that OpenAI commissioned Epic AI to produce 300 math questions for the Frontier Math benchmark. They own these and have access to the statements and solutions, except for a 50-question holdout set. So this is common with tests like this, with the Arc AGI, where you do have problems that are kind of maybe similar to how you're trying to solve it. They show how the solution is expected to go. If you've taken an exam back in the days, you've probably had like a sample problem that showed you how to solve it. And then 50 more questions that you had to solve on your own. So the 50-question holdout set, that was the 50 questions that OpenAI, that these models have never seen, that were guaranteed to be not in the data. And here they kind of spell out how the agreement with OpenAI was reached. So I'll leave this as a link below if you want to read it. I guess it did create a lot of red flags for people that the benchmark for AI was commissioned, paid for, sponsored by an AI frontier lab. So here they kind of clarify all this stuff. And so coming back to our math meeting, so the mathematicians who participated had to sign a non-disclosure agreement and could only communicate via the messaging app signal to make sure that no data was leaked. And each problem that the O4 Mini couldn't solve would garner the mathematician who came up with it a $7,500 reward. By the end of the day, Ono was frustrated with the bot whose unexpected mathematical prowess was foiling the group's progress, right? So they couldn't win the $7,500 because, you know, O4 Mini was just a little bit too good at solving these problems. So Ono here saying, I came up with a problem which experts in my field would recognize as an open question in number theory. A good PhD-level problem. He asked O4 Mini to solve the question. Over the next 10 minutes, Ono watched in stunned silence. As the bot unfurled the solution in real time, showing its reasoning process along the way. The bot spent the first two minutes finding and mastering the related literature in the field, right? So starting with some research. Then it wrote on the screen that it wanted to try solving a simpler toy version of the question first in order to learn, right? So it's kind of building a little model or prototype to kind of figure out how things worked. And a few minutes later, it wrote that it was finally prepared to solve the more difficult problem. Five minutes after that, O4 presented a correct but sassy solution. It was starting to get really cheeky, said Ono. So he's a freelance mathematical consultant for Epic AI. And at the end it says, no citation necessary because the mystery number was computed by me. So we're making AIs that are smart and sassy apparently. Is a SaaS a good thing for an AI to have or no? I wonder. So Ono's saying, I've never seen this kind of reasoning before in models. That's what a scientist does. That's frightening. Now, interestingly, they were able to find 10 questions that the bot failed to answer. And they're comparing it to a very, very good graduate student. It was also much faster than humans, taking mere minutes to do what it would take a human expert weeks or months to complete. There was a concern that the O4 Mini's results might be trusted too much, right? Just because of how, you know, sassy and strong it is. You know, mastered proof by intimidation. It says everything with so much confidence. And we've seen that, of course, with a lot of these AIs. I mean, they tend to be good. And because they're good, we tend to trust them. And then when they say something wrong, they still say it with a lot of confidence. And people can be tricked by this. So it's something to keep in mind that you have to be careful about this. And so these questions, they've sort of categorized them as tier four, right? Kind of near the top of human ability. And after going through this, the discussions turned to tier five. Those are questions that even the best mathematicians can't solve. So, of course, what happens when we reach tier five and it's super human at solving some of these questions? The roles of mathematicians would change rapidly. And so here, Arnold was saying, I've been telling my colleagues that it's a grave mistake to say that generalized artificial intelligence will never come. That it's just a computer. I don't want to add to the hysteria, but in some ways, these large language models are already outperforming most of our best graduate students in the world. So obviously, that's a very strongly written article. Here's somebody that was at this symposium. So they have a few clarifications, I guess, that are important to keep in mind. And a lot of what he's saying is not contradicting the article. So, for example, he says, to our surprise, the O4 Mini High. Okay, so it looks like he's actually saying that they've used the O4 Mini High. The article stated O4 Mini. So here on these benchmarks, it's the O4 Mini High. So they're using it at the high thinking mode. So you can see here it's slightly better than the O3, specifically when we're looking at the AIME, which is competition math problems. So he's saying this was able to solve the majority of the problems, to his surprise. While the reasoning was occasionally incorrect, it still managed to arrive at the correct numerical answers. Now, this is a little bit of a problem. Obviously, when we're doing a reinforcement learning, we're sort of checking the model's answer. We're giving it, you know, a high five if it gets the correct answer. Verifying the actual reasoning to make sure it's correct is a lot more difficult. It's a lot more resource intensive. So we've seen examples where faulty reasoning leads to the correct answer, right? And that behavior gets a plus one. It gets positively reinforced. So here he's saying the reasoning was occasionally incorrect and it arrived at the correct numerical answer. Obviously, that's a problem. It still sounds like an open problem. Hopefully, the occurrence of that is going down. But that's something that's kind of hard to prevent, I think. Just if we're doing RL on the final output and not on the actual reasoning, this sounds like it's going to happen. And I don't think we have a great answer to this quite yet, like how to fix it. And as they're saying here, AI is surprisingly effective at finding, referencing, and applying those results from, you know, including recent research results. So this person adjusted their strategy. They took a math paper, extracted some intermediate theorems, created a problem that required synthesizing those results into a computational method. And as expected, the AI struggled. It couldn't connect the intermediate steps or reason through the chain of logic effectively. So I think the article and this person, they're saying roughly the same thing. The article does seem a little bit more brosy. But they do mention, you know, 10 problems that it failed on. This person found a specific weakness, right, by taking two different things. Like it couldn't effectively synthesize them into that new computational method. Now, I'd be curious how often it would fail to do that, right? Is it every time? Is it it fails to do something like this one times in 10? But his takeaways were that, you know, AI has improved dramatically over the past two years. Current LLMs still rely heavily on pattern matching with limited deep reasoning. And they're not yet capable of generating new mathematical results. But they excel at gathering relevant literature and drafting initial solutions. Human oversight remains essential, especially for verification and synthesis. His prediction is that in the next one or two years, we'll see AI assist mathematicians in discovering new theories and solving open problems. As Terence Tao recently did with DeepMind, I'm assuming he's talking about Alpha Evolve, which I was going to talk about next. Soon after, AI will begin to collaborate and eventually work independently to push the frontiers of mathematics and by extension, every other scientific field. So as far as I can tell, they're using kind of like the normal chatbot mode, right? So they put in their thing. The chatbot thinks for whatever time it takes, right? On that high budget thinking. And then it gives the answer. Now in Google's Alpha Evolve, they reached a lot of advanced algorithms and brand new things. I think most people would argue. So Google has deployed the algorithms discovered by Alpha Evolve across a lot of their ecosystem, right? Data centers, hardware, software. At the core of this was the Gemini large language model. And it also was able to like improve, optimize its own training. So it's almost kind of starting to do recursive self-improvement. We're at sort of maybe the beginning of that stage. Another thing it came up with is how to orchestrate Google's vast data centers more efficiently. It's been in production for over a year. So the solution it came up with has been running for over a year. Improved some of the hardware on which AI runs. And here's kind of what Alpha Evolve looks like, right? So here you have the large language model, right? So this is the thing that produces the outputs. And around it, you have scaffolding, right? Some code and tools that it uses. And you have human oversight, the scientist slash engineer, right? So it comes up with prompts and the evaluation code that evaluates every output by the model. We also give it some initial programs and kind of like our initial database of stuff that we've learned up to that point. And so this thing can generate hundreds of outputs, maybe thousands, maybe more as much compute as you're willing to pay for. It's going to keep producing various solutions. Those solutions get evaluated and similar. It's kind of like an evolutionary tree. The promising solutions get worked on more. So it's kind of like a offspring or lineage that keeps getting improved. Here's a great representation. So this is the Darwin Godel machine, which uses something very, very similar. So this is for creating a better coding agent. So it's a kind of a self-improving coding agent. So we start at zero and then it tries like, here's one theory of how we can improve it. One, two, three, four, right? So it's sort of different. It's brainstorming different ideas. Those ideas get tested on a benchmark. If they improve, they get marked as promising. If they perform poorly, the red ones means that they didn't work that well. And those lineages get cut. So similar to evolution, things that don't, that are not adapted, kind of don't survive, don't pass on their genetic information. This is kind of similar. And so here, as you can see here, this lineage is the one that eventually through many, many iterations, 80 in total, it leads to the best coding agent here. And it works extremely well because this purple line, this is an existing human coding agent. So somebody, the human being sat down and coded up a bunch of stuff to make that agent better at coding. So that's the purple line. And this blue line here is this Darwin Godel machine. So as you can see here, it starts out worse than the human made one. But as it generates a bunch of little potential solutions as they get tested, eventually it figures one out and it jumps up. And then another, another, you know, generates a bunch more that maybe don't work. But eventually it keeps getting better and better. So this is an example of an autonomously improving coding agent that gets better than the human baseline. And then with this other human baseline, this was the best state of the art coding agent, this red line. It starts much worse and slowly approaches it. So it's important to understand here that this sort of evolutionary search produces often some incredible results. But the large language models are allowed, in this case, 80 iterations. And then Alpha Evolve also had, I don't remember how many iterations per thing, but it could have been many. As long as they're willing to spend money on compute on having it do that kind of open-ended exploration. Like the more you do, the more likely you'll find an answer. So the idea in a nutshell is having it do, try a lot of different stuff, test that stuff to see how well it works. And then if a certain avenue seems promising, you build on top of it. With this symposium that we're talking about, they're not doing that. I think they're just probably doing one output. There's no feedback. There's no evolutionary search. There's none of that, really. So that would be equivalent to, you know, here we have an output one thing, right? That thing doesn't work. And we say, oh, it failed, right? That's it. Whereas in reality, given, you know, 80 iterations, giving some way of testing this and doing that evolutionary search, those results could have been staggering. That's just my theory. But I feel like there's probably a way to combine the findings of Alpha Evolve and the Darwin machine, kind of like how they create those things and combining it with what they're doing here. Because, again, these models can come up with a thousand outputs just as easily as they can with one. It costs a little bit more. It takes a little bit more time. But it's not going to get bored or not want to do it. It'll generate as many as you ask for. And if you're just able to create some sort of system that gives it feedback and has it kind of continuously improve, then it seems like the results could be much better. So here he's saying human oversight remains essential, especially for verification and synthesis. So things like Alpha Evolve do automate some of it. You still need that human loop kind of overseeing it. But if you're able to automate verification and some of that synthesis, then this thing gets kind of supercharged. I'll also link this down below. So this is Gary Marcus. So he comes in with some of the counter arguments saying, we have no idea what math specific augmentation OpenAI has done. The big question is about math relevant augmentation, which I think to be did OpenAI train it on additional data. I think this Epic AI frontier math benchmark was probably some of that data that helped these models better understand how to do the high level math. Except for the 50 question holdout set. Again, if I understand correctly, that's probably what happened. But let me know what you think about this whole thing. Are these models getting better at math or do you think it's just simple pattern matching? It's kind of an illusion. Can the results where, you know, the world's best mathematicians are stunned by its reasoning abilities, Google deep minds, AI systems getting, you know, one point away from the gold medal at the IMO. This self-improving coding agent that gets either almost as good as or far better than kind of human innovation. How far we were able to get these coding agents to go and everything that Google deep minds, alpha valve, everything that it was able to optimize. Like, do you think it's possible that this is just trickery illusion? This is a simple pattern matching. You know, here the vast data centers of Google are saving 0.7 of Google's worldwide compute resources based on this optimization that the program came up with. If you think about how giant Google's data centers are, I mean, that's a lot of money. It's been in production for over a year. So that's not only saving, what, millions, tens of millions, maybe more, right? It also means that more tasks can be completed on the same computational footprint. So that's more efficiency, less energy use, et cetera. I'll leave you with the words of Scott Aronson that I think puts it really well. So he worked for Google on their quantum supremacy project. He worked a little bit for OpenAI and just a kind of a smart guy all around. So take a listen and I'll see you in the next one. To my lifelong chagrin, people are just constantly munging these questions together, right? They're just constantly saying, hey, I will never be able to do these things because it doesn't really feel or it's just simulating it. It doesn't really have that inside. And then once it does do that task, then they just shift to a different thing that it will never do. And then it does that thing and so forth. OK, so there is I was trying to come up with a name for it. I'm going to call it the religion of justism. OK, so there's like there's this whole sequence of deflationary claims, right? Each person who makes them thinks that they're like the first one, right? And they there's like I've seen like 500 different variants of this now, right? Chat GBT, it doesn't matter how impressive it looks because it is just a stochastic parrot. It is just a next token predictor. It is just a function approximator. It is just a gargantuan autocomplete, right? And what these people never do, what it never occurs to them to do is to ask the next question. What are you just a, right? Right. Right. Aren't you just a bundle of neurons and synapses, right? I mean, like we could take that deflationary reductionistic stance about you also, right? Or if not, then we have to give some principle that separates the one from the other, right? You know, it is our burden to give that principle. So the way that someone was putting it on my blog was, OK, they gave this giant litany. Look, GPT does not interpret sentences. It seems to interpret them. It does not learn. It seems to learn. It does not judge moral questions. It seems to judge moral questions. And so I just responded to this. I said, that's great. And it won't change civilization. It will seem to change it. OK. Here we go. Hey. And so I'm just a kid. I'm just a kid. I'm just a kid. And so I'm going to move thisStrose.