Open transcript
So the Arc AGI benchmark is back, including the new Arc Prize for 2025. It's back and better than ever with a completely new set of questions that stump even the smartest AI models, but humans can pass. If you notice here, the 03 low spends like $200 per task and still achieves a score less than 5%. The new test is also a lot more resistant to test time compute. So just throwing more money, more compute at it is not going to solve it. So as you can see here, we saw sort of an improvement on the score of the Arc AGI, including the 03 low. And basically, as long as we spent more money doing test time compute, it would score higher. It would be a lot more expensive, but it would score higher. So this Arc AGI 2 grand prize is kind of going to be located here. Purely scaling of the cost it takes to run these things isn't going to get it there anymore. So the grand prize will go to somebody that's going to get 85% accuracy at around a 42 cents per task efficiency ratio. So you can't spend that much money in compute per task. And on this updated set of questions, and as you'll see also, they threw a few new sort of dimensions in there to make the test harder. Currently, the base LMs, the non-reasoning models, they score a zero. The reasoning systems score less than 4%. So as I say here, this is an unsaturated frontier AGI benchmark. Now, interestingly, every Arc AGI 2 task is solved by at least two humans quickly and easily. We know this because we tested 400 people live. So interestingly, they're specifically seemingly, they kind of target these things that these reasoning models and LMs are bad at. But it's very easy for humans. I'll show you exactly what I'm talking about in just a second. So the point of Arc AGI 2, it's not about these LMs showing superhuman skills. It's about exposing the thing that's missing in AI. It's an efficient acquisition of new skills. So it's not data memorization or some sort of a pattern recognition. The question is, can you sort of acquire new skills as you're going through it? It challenges capabilities like symbolic interpretation, compositional reasoning, and contextual rule application. So with symbolic interpretations is understanding that these figures can mean certain things beyond just what they look like. They're a symbol. So given this example, like this converts into this, therefore, what does this convert into on this grid? We also have AI reasoning systems that struggle with tasks requiring simultaneous application of rules or multiple rules that interact with one another. There's one problem that I solved that really illustrated this in a really great way. And the idea of struggling with contextual rule application. So basically where rules must be applied differently based on the context. All right. So this is the daily puzzle for the new Arc AGI prize. Let's try it out. Oh, it's a timed example. All right. So here's our example. So it looks like it's shifted three to the right or rather green is shifted one, two to the right. Let's say the bottom one is shifted two to the right. So step one, I have to make it into whatever this is. So 13 by 14. Okay. Let's resize that. All right. So first things first, here are the blue and the red is shifted over by one, two. So one, two, and it's four. So I'm assuming the bottom one is the one that gets shifted. So it goes one, two like that. Yes. And this confirms it. Okay. Oh, I should have looked at all the different examples. Okay. So this one's showing, oh no, it's the same thing. They're all moving over by two. I think mine is right. Let's test it. Wrong. All right. So the top has moved one to the left and the bottom one's moved right one to the left. Okay. Gotcha. All right. I think this should do it. I'm guessing they want me to color this in. Okay. Select. All right. So let me just fill these out. So that should do it. Boom. I am AGI. All right. Very good. All right. And here I'm going to try the public training set V2. So this one's going to be easy. So the examples are from this to this. So as far as I can tell, first of all, we go from a two by two to a six by six. And so this just kind of gets copied here and here and here. Then on the next row, you can kind of think of it as flipping it vertically and then kind of horizontally. So basically, or just flipping it diagonally and then going boom, boom, boom. And then doing that again here. Boom, boom, boom. And let's confirm. So the pattern seems to match here in the second portion. Nope. I missed something. Let's see. It looks like the bottom left shifts to the bottom right and top right shifts to top left. Yeah. So the pattern holds and then it does that again. Okay. I get it. First and foremost, we have to configure our grid. So this one's going to be a six by six. All right. So let's start with the greens. Boom, boom, boom. And the second row, it's going to shift to the right like that. And then it's going to revert here. And then we have this red color that just stays here, shifts and stays. Then we have the orange, then shifts and then stays. And then we got this blue and like so. Did I get it? Boom. Try the next puzzle. All right. Let's bump it up to, let's go for the hard. What does the hard look like? All right. So we got this going on here. So this translates into this. The grid remains the same. Okay. So that's, that's very, very good. Okay. So in all of these sort of red line that runs vertically, it's some sort of a like translation. So you take here, this stays as is. And then based on this pattern, it translates into what happens sort of on this side of the red line. So here, one gray translates into a row of grays. Here, two blue translates into a blue every other time. All right. Okay. So that kind of makes sense. And three oranges translates into three every three or an orange every three. Okay. So this one kind of like this example really spells it out. But what happens when you have a multicolored ones? So here, four blues, it translates into blue every four, five greens translates into a green every five. All right. So we know what this looks like. Like the red one's going to be red every five. So here, and then every five like that blue is going to be every two. So these make sense to me. If it's one color, that makes sense how that works. So what does that mean for, for example, here? So if I'm reading this correctly, so basically both obey the same rules, but you apply like, for example, this teal color or whatever, I don't know, like a light blue, right? So because it's at one, it just goes and it covers everything. So this one doesn't get applied. So maybe it's a thing where you sort of like, this would be this purple color every four, but then this is every single one. Okay. So I'm going to assume that's what it is. So we need a, so this is a 20 by 10. So we're going to need a 20 by 10 grid, 20 by 10. Oh, and just to be difficult, they threw in a yellow color here. So there's five spaces. Then we have our yellow color that doesn't go all the way to the bottom. So the gray one, this means every four. So it goes like this, then like that. The blue one means every two. So kind of like how they show it here, we go like that. And then the red one, every single one. That's great. So then the orange, the orange, since it's one, I mean, it would kind of be like this, right? So it's like that, but then we're going to cover the purple over it. So this one means it's every two. So kind of like that. And these are purple. And then we have this one was two purple and like that. So two means it goes every two. And this means it goes every three. So like every three like this, basically. So it goes one, skip two, one like that. So one, skip two, one, skip two. All right. This is it. Let's go. Oh yeah. Winner, winner, chicken dinner. All right. So they seem pretty straightforward. They're a little bit complicated at first, but if you kind of just sit there for a little bit, it kind of like you start noticing the patterns. And yeah, it's one of those things where humans would be pretty good at this. So it sounds like humans beat all of them, or at least capable of beating all of them, but large language models might struggle on figuring out how to do it. So I'm not going to go through all these, obviously, but just from the few that I tried, it seems pretty good. And as they've mentioned earlier, efficiency will also count here. So by intelligence, they're no longer just looking at, does the capability, the accuracy, they're also looking at efficiency. They can no longer report performance as a single metric. So this leaderboard will also now track the cost of the performance. So the reasoning models were sort of adaptive to the test because you can crank up how much you would think about any given problem and thereby improve the score. This new test is resistant to that sort of ability. By the way, let me know in the comments what you think about this. Is this, you know, quote unquote fair? Is this good to limit how much time you can spend thinking? Obviously, as we improve the efficiency of these models, it'll cost less, but is that a good metric? The other interesting thing is it's no longer just about climbing the leaderboard and trying to get the highest score possible. They've adjusted the scoring to incentivize conceptual breakthroughs. There's a $75,000 prize for the most significant conceptual contribution, a $50,000 prize for the highest score. The grand prize has been increased to $700,000. So definitely recommend that you go and click on play. It gives you a daily puzzle, which by the way, is timed and that does get reported along your score. I think you'll be happy to know that I'm still more generally intelligent than an AI and I have my certificate here to prove it. They cranked up the difficulty just a little bit. Again, not in terms of like you need some massive processing power. You just need to kind of pay attention and try to figure out what the patterns are, what the rules are. And here's our current leaderboard. So as you can see here, the human panel, well, they got 100%. So out of the people that are taking, I think each question, at least two people solve it. So I think they're using kind of a group of people and at least if you solve it, then it counts as being solved. So all the questions that are included were solved by at least two people in that group. Then we have the O3 low chain of thought and a synthesis, right? Comes in at 4% and costing $200 per task. And it kind of drops off from there 3% for the O1 high. Architects, which received the ARC price from 2024, gets a 2.5%. But that's the first model here that actually comes under the cost per task. DeepSeq R1, the Chinese open source model, noticeably is also here kind of high up there. 1.3% and just $0.08 per task. So definitely a good showing. There's some rumors about the new upcoming DeepSeq reasoning model doing very well on the RKGI. I can't, you know, verify those rumors. So they're just floating out there. I don't know what the source is. So maybe, maybe it's true. Maybe it's not. We'll see. We are still expecting DeepSeq. I'm guessing they're going to call it R2, the next iteration of this one, to be dropped pretty soon. So exciting to see a fresh unsaturated benchmark, as they put it. Again, the best AI is 4% and way above in the cost per task. The first opening AI model that will qualify is at only 1.7%, right? So we have a lot of room for them to improve, to go, to climb this leaderboard, to get to the 85% needed to qualify for the grand prize. Now, this is from the manifold market. So we're able to kind of place bets on various outcomes here. They're asking, will the ARC AGI grand prize, so that's the first version, 2024 data set, is anybody going to claim it by the end of this year, 2025? So the market thinks there's a 27% chance. It looks like there's another betting market here. If anyone will score 70 plus percent on the ARC AGI 2 within three months of its release, people are thinking there's an 8% chance. What do you think the chances are? What would you guess? But here's some food for thought. So this is Isaac Liao, machine learning PhD, previously computer science and physics at MIT, international physics Olympiad 2019 silver. Smart guy, I would say. I think that's safe to say. I'm going to give him a follow. But, so he reposted the ARC AGI prize too, but if you keep scrolling, he's got this. Introducing ARC AGI without pre-training. No pre-training, no data sets, just pure inference, time, gradient, descent on the target ARC AGI puzzle itself, solving 20% of the evaluation set. So interestingly, we have a lot of different approaches, people trying out new stuff. The organizers of the ARC prize themselves, they're not really doing this to, you know, have everybody lose. They are looking for somebody to actually come up with the correct solution. Francois Chalet is giving some sort of a hints about what he thinks might be the right approach. But of course, the whole point of this is for everybody all over the world, everybody that wants to compete to try this out, maybe come up with something new and innovative. It has to be open source. So whatever new things we discover from it, everybody else will be able to kind of learn about those and try them out. So let me know what you think about this whole thing. Do you agree with them limiting how much a compute sort of you can spend on it, what the budget is? This seemed kind of low, 42 cents per task. Do you agree that that should be a restriction? I mean, they still had their $10,000 sort of total compute that was previously. But let me know what you think about that new restriction. What do you think about the test in general? Certainly, they seem well organized to kind of push innovation forward for people to come up with new ideas, to be incentivized, to put them out there and to hopefully, you know, kind of an open source, a group effort to try to push this thing forward. I got so excited about being generally more intelligent than an LLM. But I think I jumped the gun there. Because at the end, I said, are you starter than an LLM? Which of course should say, are you smarter than an LLM? So I think I won the battle, but I lost the war. But check it out for yourself. Play the game. Tell me what you think about it. If you made this far, thank you so much for watching and listening, and I'll see you next time.