← Back to video archive

Airdroplet AI summary

Apple's STUNNING Discovery "LLMs CAN'T REASON" | AGI cancelled

June 9, 2025Wes RothAI score 10068,180 views

Watch original on YouTube ↗

AI-generated summary

This video dives into Apple's recent research paper, "The Illusion of Thinking," which suggests that large language models (LLMs) can't truly reason and that their apparent thinking abilities are just an illusion. The presenter, Wes Roth, unpacks Apple's claims, critiques their methodology, and offers a counter-perspective, arguing that the models'

Video transcript

Open transcript
Steve Jobs said this, we humans are tool builders. A human on a bicycle blew the condor away. And that's what a computer is to me, a bicycle for our minds. Now we have AI, a tool that builds tools. I bet Apple is all over this. So Apple publishes a paper called The Illusion of Thinking, where they test reasoning models, so large language models with reasoning abilities, to see how well they perform certain tasks. A lot of people have been posting that now we're realizing that large language models can't reason, that it's all an illusion. So naturally, I wanted to take a look at this paper because I was very curious to see what is it that they found out that they can definitively say that these models can't reason, that it's just an illusion. Now, interestingly, this is not the first paper out of Apple. They kind of have this cynical approach towards large language models. There was another one here saying, understanding the limitations of mathematical reasoning in large language models. Now it's realizing the limitations of reasoning through problem complexity. Now, Apple doesn't have any reasoning models that I'm aware of. If they do, it's nowhere on any leaderboard. Apple has, I think we can all agree, the worst AI products. So it is a little bit strange that they're just publishing research on what everybody else is doing, kind of saying, well, here's why it doesn't work. As Andrea White here puts it, no idea what their strategy is here. Me neither. I just don't really get why they're doing this, but what are the cases? Let's take a look at the paper. And so really fast, they're saying how the latest frontier language models introduced large reasoning models. So we have LLMs where you ask a question and it immediately sort of spits out the answer. And more recently, we have the introduction of what some are referring to as large reasoning models. So it's the ones where it says, I'm thinking about it for 20 seconds and it has a chain of thought reasoning that has been shown to improve some of its answers, especially on certain tasks where that reasoning helps. Usually it's math, coding, basically anywhere where thinking through something before answering might help out. And they're saying that current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy. And as you'll notice through the paper, they do mention this idea that these benchmarks aren't the best way of testing these models because we're not sure if this was in their training data and they've just memorized those answers or they're actually thinking through it and coming up with new insights after they've reasoned through it, et cetera. And this is certainly true. We've seen certain AI labs kind of game these benchmarks by training their models on data and they tend to do really well in those benchmarks. But then when you try to get them to do real world tasks, they don't do as well as expected. That's why a lot of people kind of take these benchmarks with a grain of salt. They can be good at a first glance, like a first snapshot into how well a model is expected to perform. But at the end of the day, you've got to test it out yourself for your own use case to see if it's good or not. Before we continue, just a really quick illustration of what we're talking about between kind of easy answers where you can kind of give that quick response and it might be the correct one versus something that's a little bit more reasoning intensive. So if I say, what is two plus two, right? You might say four. You don't have to think about it. You don't have to expend a lot of mental energy. The answer is right there. You don't have to really think through it. Now, if I say, what is nine times a six? Now, it might take you a little bit more effort to think through it, right? Maybe you have it memorized potentially, right? If you did recently or just tend to have these things memorized, you might know the answer right away, or you might have to do the math in your head, or you might remember some trick to doing that calculation. But the point usually is going to take you just a little bit of time to think through it and come up with the answer. But, and this is the last question I'm going to ask you, how many prime numbers are there between one and 15 million? Okay. Think about it really hard. Do you know what the answer is? Now, for most of the people, we might quickly think like, is there some shortcut to doing that problem? And if we can't think of one, it's unlikely that you're going to start expending tons of mental effort into trying to solve this problem right now as you're watching this video. So you might have expended a little bit of effort here more for this problem, but for this problem, you probably spent less effort. You realized that it was very complicated and chose not to engage with it. Point being that the amount of time you're willing to spend on the problem doesn't just increase the harder problem gets. At some point, you kind of estimated that's a little bit too difficult for you to bother with, and you just don't think about it anymore. Keep that in mind. Now, what were the main findings in this paper? There was three. One, low complexity tasks where standard models surprisingly outperforms large reasoning models. So like two plus two, you might quickly know the answer. You don't need to think about it. There are probably even scenarios where overthinking could, you know, reduce the result. Number two, medium complexity tasks where additional thinking in the large reasoning models demonstrates advantage, right? So kind of those a little bit more complicated things where you have to think through it step by step. That's of course where the ability to think through it step by step really shines, right? So that's where the results improve. And three, high complexity tasks where both models experience complete collapse. So in other words, if the problem is very complicated, very complex, then both types of models just collapse. They can't continue. And these models are tested on the Tower of Hanoi. Checkers jumping, river crossing, blocks, world, etc. And here kind of the main charts, as you can see, for, you know, these sort of easy problems, the thinking and the non-thinking models are kind of similar. In the blue, so kind of like for the kind of medium difficulty challenges, right, the thinking models get a distinct advantage. So when the ability to think longer before giving an answer in certain tasks that plays a role and that creates a big advantage for these models kind of makes sense. Right. Again, two plus two, you know the answer immediately. You can just spit it out nine times six or whatever we used in this example. You know, unless you have it memorized, that may take you a few seconds. You'll still get the answer, but you just need a few seconds to kind of process through it. And here across the sort of the harder tasks, that's where both models kind of completely collapse and get zeros. They're unable to continue thinking through the steps and coming up with correct answers. As you can see, they did it with CLODE. They did it with DeepSeq R1. And they also used O3 Mini, the medium and high configurations on some of these tests. And in the conclusion, they said that these models failed to develop generalized reasoning capabilities beyond a certain complexity threshold. And so this is probably where the people get this idea that reasoning doesn't work. It's not real, that it's an illusion. And tons of people are kind of writing about saying, well, you know, Apple drops the mic. They're saying that large language models don't think. They assimilate thinking. And you need to understand this, that LMs are guessing machines. Here's another paper saying LMs often assimilate logic without truly understanding it. There's many, many more like this. But let's look at some of the counterpoints to this study. One of the best explanations of some of the issues with this paper, I think, goes to, I think, this blog post, Sean Godeke. I'll link him down below. Unfortunately, he's not tweeting since a number of years ago. I don't know why. Sean, if you're listening, get on Twitter. We'd love to have you. Now he's saying, I do not believe that AI language models are on the path to super intelligence, but I still don't like this paper very much. And again, this is kind of my take as well. I see certain issues with it. I'm not trying to be the champion of large language models and defend their honor or anything like that. I just, there's some questions that I have about this paper that I think need to be answered. First problem with the paper is they kind of say that the coding and math benchmarks are bad because those problems can exist in the training data and they choose Tower of Hanoi. But as Sean says here, Tower of Hanoi is an even worse test case for reasoning than math and coding. If you're worried that math and coding benchmarks suffer from contamination, as in that data being part of the training data with the large language models, why would you pick well-known puzzles for which we know the solutions exist in the training data? He links to Google results for 10-disc Tower of Hanoi solutions that shows, first of all, the Google AI overview, tons of videos, tons of results. I mean, there's going to be a lot more examples of how to solve this on the internet than some obscure benchmark, right? There's not going to be, you know, a million pages about some coding AI benchmark. There is like a million pages about how to solve the 10-disc Tower of Hanoi puzzle. And so he continues, because of this, I'm puzzled by the paper's surprise that giving the models the algorithm didn't help. Because at some point they gave the model the algorithm to see if that improves its abilities, right? So the Tower of Hanoi algorithm appears over and over in the model training data. Of course, giving the algorithm doesn't help much. The model already knows what that is, right? So we're not giving it new information with which to be better. Like it already knows it. It's like if you were ever like playing a video game and somebody's like, have you tried not dying as much or something like that? You're like, no, I am aware of the goal of the game is just, I'm still struggling with it. You're saying, you know, try not to suck doesn't help. Also, if reasoning models have been deliberately trained on math and coding, not on puzzles, right? So with reinforcement learning, we know that a lot of the companies, they want them to do well at math and coding. That's the focus. Are puzzles a fair proxy for reasoning skills? Maybe, maybe not, you know, certainly I would bet more on coding skills and mathematical skills over puzzle skills, right? He gives the example of like, it's kind of like saying, well, these models haven't gotten better at writing patriarchal sonnets since GPT 3.5. So I don't think any real progress has been made. Again, picking these weird out of the way puzzles that doesn't necessarily show anything. And this might be a case of the streetlight effect. So basically we study things where it's easier to observe what's happening. It's like if you lose your car keys somewhere outside of the mall, and you only look for them under the streetlight because it's dark everywhere else. So you're not looking for them. You're just looking underneath that shining light. They're not more likely to be where you can see them. They might be somewhere in the dark. So just looking underneath the light doesn't really help. So a number of people, including Sean, they've actually tested some of these prompts with DeepSeek R1. So they gave it to the Tower of Hanoi puzzle with 10 disks. The model does some basic maths and realized that, you know, it's a lot of moves. So generating all those moves manually is impossible, right? Kind of thinks through it. Remember that example I gave you of figuring out how many prime numbers there's between 1 and 15 million. You probably maybe like thought, is there some shortcut? No. So you would actually have to do the math to figure it out. You have to, you know, meticulously count them all, and you gave up, right? So guess what these reasoning models do? The same exact thing, right? Note the model immediately decides that generating all those moves manually is impossible because it would require tracking over a thousand moves. So it spins around trying to find a shortcut and fails. So the key insight here is that past a certain complexity threshold, the model decides that there's too many steps to reason through and starts hunting for clever shortcuts, right? So past eight or nine disks, the skill being investigated silently changes from can the model reason through the Tower of Hanoi sequence to can the model come up with a generalized Tower of Hanoi solution that skips having to reason through the sequence, right? So notice here in the paper, this is the 10, right? So right here, there's that line where it gets too complicated. And as you can see here, like the results are close to zero. And so the researchers said the accuracy of this model past, you know, 10 or whatever is zero, right? They keep marking it wrong. Here's somebody on X that actually also replicated this experiment, saying a few more observations after replicating the Tower of Hanoi game with their exact prompts. So again, using the exact prompts from that Apple paper. So this person notes, you know, with the context window, the output limit for these models are, and saying all models will have zero accuracy with more than 13 disks simply because they cannot output that much. The max solvable sizes without any room for reasoning is, you know, 12 disks on DeepSeek and 13 disks on a Sonnet 3.7 and O3 Mini. Notice they're not testing the million token context window Gemini. And if you actually look at the output of the models, you will see that they don't even reason about the problem if it gets too large. Due to the large number of moves, I'll explain the solution approach rather than listing, you know, the 32,000 moves individually. So I hope what's happening here makes sense to you, right? Just because the model thinks through the complexity of the problem goes, well, there's no way I'm going to list all of the steps and doesn't even try. That's not the same as it's getting it wrong. At the beginning of the video, when I asked you 2 plus 2, you might have humored me and said 4, right? You kind of instinctively knew the answer when I said 9 times 6. Maybe you took a second or two to think about it. Maybe you remember some shortcut. Maybe you even got the answer. And then when I gave you this insanely complicated thing to do, you might have thought about it for a few seconds, realized that there's no simple way to solve it, and just gave up. If you did, that means you can't reason, apparently. You're incapable of reasoning. Your reasoning abilities are an illusion, I'm sorry to say. Or is this a very human reasoning ability to think through and say, I'm not going to do this thing because there's not a simple way of doing it. I can't complete this task. So let me think about doing a shortcut, right? It's a more human way of approaching it. These models emulate how we reason about things. By the way, if you asked me to solve the 10-disc Tower of Lannoy problem, you know, do you think that I would like just sit there and do it? Take the time and meticulously write out the however many thousands of moves there are? No, you know what I would do? Here I asked Gemini 2.5 Pro to create some code to solve the Tower of Lannoy puzzle with 10 discs. Here's the result. Let's see what it did. So you can change the number of discs here. Let's set it to 10 and start. Let's see if it can in fact figure it out. As you can see, it knows that the minimum amount of moves it would need is 1023. It's at 600 right now. So let's see. Let me bump up the speed a little bit because it's like animating everything. All right, there. The Tower of Lannoy solver. It solved it in 1023 moves. But wait, I thought reasoning was an illusion. So how come it's able to build a tool that solves the problem that we gave it? How come it's able to understand the fact that for these very complicated problems where you have to list over a thousand moves that it just might not be viable and it has to think of a different approach? That seems reasonable, doesn't it? For lack of a better word. But let me know what you think about this. Do you think Apple is onto something and that these models are just sort of have an illusion of thinking and reasoning? Again, reasoning and thinking, we don't have clear definitions for those words, especially in terms of like how do they relate to AI. Better question is, does it seem like it's doing something that's similar to how humans think and reason? Maybe there's some parallels there. Maybe it's some imitation of it that seems to work, that seems to achieve some of the same results. Or do you think that Apple is falling behind in their AI development and that Steve Jobs is turning over in his grave? Let me know in the comments. My name is Wes Roth. Thank you so much for watching.