← Back to video archive

Airdroplet AI summary

ARC AGI 2 The $1,000,000 AGI Prize

March 25, 2025Wes RothAI score 10025,077 views

Watch original on YouTube ↗

AI-generated summary

Okay, let's dive into what's happening with the new ARC AGI 2 prize and benchmark!

Here’s a quick rundown: The ARC AGI benchmark is back with version 2, featuring a completely new set of problems designed to expose the current limitations of even the best AI models, particularly their lack of efficient reasoning and ability to learn new skills on the fly, unlike humans who find these tasks easy. There's a big prize pool, including a $700,000 grand prize, for someone who can achieve 85% accuracy efficiently, pushing AI development towards true, human-like intelligence rather than just scaling up compute.

Here are the key points from the video:

  • The ARC AGI benchmark has been updated to ARC AGI 2 with a new set of questions that are hard for AI but easy for humans.
  • The new test is designed to be "unsaturated," meaning current AI models score very low, highlighting that we are still far from AGI on this specific measure.
  • A major change is that the new test is resistant to "test time compute," which means you can't just throw more money or processing power at a problem to get a higher score anymore.
  • In the previous version, increasing the cost of running a model per task often led to higher scores, but that won't work for the grand prize in ARC AGI 2.
  • The grand prize requires achieving 85% accuracy on the new questions at an efficiency ratio of around 42 cents per task.
  • Current base LLMs score zero on the new test, while existing reasoning systems score less than 4%. That feels pretty low compared to human capability.
  • Intriguingly, every single ARC AGI 2 task has been solved quickly and easily by at least two humans tested live, suggesting they target something specific that humans are good at and AI isn't.
  • The goal of ARC AGI 2 is not to showcase AI's superhuman abilities but to identify what's fundamentally missing for AGI, focusing on the efficient acquisition of new skills.
  • This is not about data memorization or just pattern recognition; it's about understanding and applying rules in novel situations.
  • The test challenges capabilities like symbolic interpretation (understanding that shapes represent concepts beyond their appearance), compositional reasoning (applying multiple rules that interact simultaneously), and contextual rule application (changing how rules are applied based on the specific situation).
  • Demonstrating solving a few puzzles from the test showed that while initially complex, humans can figure out the underlying patterns and rules through reasoning, even if they make mistakes at first. It feels like solving logic puzzles.
  • Efficiency is now a tracked metric on the leaderboard alongside accuracy, meaning performance is no longer reported as a single score but also considers the cost per task.
  • This change aims to incentivize finding more computationally efficient methods for achieving intelligence. There's a question raised about whether limiting compute time is a "fair" way to measure intelligence, but it clearly pushes towards efficiency.
  • The scoring structure has been adjusted to incentivize conceptual breakthroughs more than just climbing the leaderboard with incremental improvements.
  • There's a $75,000 prize for the most significant conceptual contribution, $50,000 for the highest score, and the grand prize is now $700,000.
  • The current leaderboard shows humans at 100% (since all included tasks were solved by at least two people), while the best AI models score very low (e.g., O3 low at 4% but costing $200/task, Architects at 2.5% and lower cost, DeepSeq R1 at 1.3% and just $0.08/task).
  • The best models are still far below the 85% accuracy needed for the grand prize, showing lots of room for improvement.
  • There are betting markets on whether the previous ARC AGI (2024 data) grand prize will be claimed by the end of 2025 (27% chance) and whether anyone will score 70%+ on the new ARC AGI 2 within three months (8% chance). These odds suggest experts are skeptical of rapid progress on this specific benchmark.
  • New approaches are being explored, such as one project (by Isaac Liao) attempting to solve 20% of the evaluation set without any pre-training or datasets, just using inference time gradient descent on the puzzle itself. This highlights that researchers are trying fundamentally different ways to tackle these problems.
  • A key requirement for the ARC prize submissions is that they must be open source. This is great because any successful approaches or discoveries will be shared with the community, accelerating collective progress towards AGI.
  • It's encouraged to try out the daily puzzles yourself on the ARC prize website (arcprize.org) to get a feel for the types of problems and see how you stack up against current AI (spoiler: you're likely smarter at these specific tasks!).

Overall, the ARC AGI 2 benchmark is a fascinating attempt to define and measure a specific type of flexible, efficient intelligence that current AI models lack, pushing the field beyond brute-force scaling towards genuine reasoning and skill acquisition.

Video transcript

Open transcript
So the Arc AGI benchmark is back, including the new Arc Prize for 2025. It's back and better than ever with a completely new set of questions that stump even the smartest AI models, but humans can pass. If you notice here, the 03 low spends like $200 per task and still achieves a score less than 5%. The new test is also a lot more resistant to test time compute. So just throwing more money, more compute at it is not going to solve it. So as you can see here, we saw sort of an improvement on the score of the Arc AGI, including the 03 low. And basically, as long as we spent more money doing test time compute, it would score higher. It would be a lot more expensive, but it would score higher. So this Arc AGI 2 grand prize is kind of going to be located here. Purely scaling of the cost it takes to run these things isn't going to get it there anymore. So the grand prize will go to somebody that's going to get 85% accuracy at around a 42 cents per task efficiency ratio. So you can't spend that much money in compute per task. And on this updated set of questions, and as you'll see also, they threw a few new sort of dimensions in there to make the test harder. Currently, the base LMs, the non-reasoning models, they score a zero. The reasoning systems score less than 4%. So as I say here, this is an unsaturated frontier AGI benchmark. Now, interestingly, every Arc AGI 2 task is solved by at least two humans quickly and easily. We know this because we tested 400 people live. So interestingly, they're specifically seemingly, they kind of target these things that these reasoning models and LMs are bad at. But it's very easy for humans. I'll show you exactly what I'm talking about in just a second. So the point of Arc AGI 2, it's not about these LMs showing superhuman skills. It's about exposing the thing that's missing in AI. It's an efficient acquisition of new skills. So it's not data memorization or some sort of a pattern recognition. The question is, can you sort of acquire new skills as you're going through it? It challenges capabilities like symbolic interpretation, compositional reasoning, and contextual rule application. So with symbolic interpretations is understanding that these figures can mean certain things beyond just what they look like. They're a symbol. So given this example, like this converts into this, therefore, what does this convert into on this grid? We also have AI reasoning systems that struggle with tasks requiring simultaneous application of rules or multiple rules that interact with one another. There's one problem that I solved that really illustrated this in a really great way. And the idea of struggling with contextual rule application. So basically where rules must be applied differently based on the context. All right. So this is the daily puzzle for the new Arc AGI prize. Let's try it out. Oh, it's a timed example. All right. So here's our example. So it looks like it's shifted three to the right or rather green is shifted one, two to the right. Let's say the bottom one is shifted two to the right. So step one, I have to make it into whatever this is. So 13 by 14. Okay. Let's resize that. All right. So first things first, here are the blue and the red is shifted over by one, two. So one, two, and it's four. So I'm assuming the bottom one is the one that gets shifted. So it goes one, two like that. Yes. And this confirms it. Okay. Oh, I should have looked at all the different examples. Okay. So this one's showing, oh no, it's the same thing. They're all moving over by two. I think mine is right. Let's test it. Wrong. All right. So the top has moved one to the left and the bottom one's moved right one to the left. Okay. Gotcha. All right. I think this should do it. I'm guessing they want me to color this in. Okay. Select. All right. So let me just fill these out. So that should do it. Boom. I am AGI. All right. Very good. All right. And here I'm going to try the public training set V2. So this one's going to be easy. So the examples are from this to this. So as far as I can tell, first of all, we go from a two by two to a six by six. And so this just kind of gets copied here and here and here. Then on the next row, you can kind of think of it as flipping it vertically and then kind of horizontally. So basically, or just flipping it diagonally and then going boom, boom, boom. And then doing that again here. Boom, boom, boom. And let's confirm. So the pattern seems to match here in the second portion. Nope. I missed something. Let's see. It looks like the bottom left shifts to the bottom right and top right shifts to top left. Yeah. So the pattern holds and then it does that again. Okay. I get it. First and foremost, we have to configure our grid. So this one's going to be a six by six. All right. So let's start with the greens. Boom, boom, boom. And the second row, it's going to shift to the right like that. And then it's going to revert here. And then we have this red color that just stays here, shifts and stays. Then we have the orange, then shifts and then stays. And then we got this blue and like so. Did I get it? Boom. Try the next puzzle. All right. Let's bump it up to, let's go for the hard. What does the hard look like? All right. So we got this going on here. So this translates into this. The grid remains the same. Okay. So that's, that's very, very good. Okay. So in all of these sort of red line that runs vertically, it's some sort of a like translation. So you take here, this stays as is. And then based on this pattern, it translates into what happens sort of on this side of the red line. So here, one gray translates into a row of grays. Here, two blue translates into a blue every other time. All right. Okay. So that kind of makes sense. And three oranges translates into three every three or an orange every three. Okay. So this one kind of like this example really spells it out. But what happens when you have a multicolored ones? So here, four blues, it translates into blue every four, five greens translates into a green every five. All right. So we know what this looks like. Like the red one's going to be red every five. So here, and then every five like that blue is going to be every two. So these make sense to me. If it's one color, that makes sense how that works. So what does that mean for, for example, here? So if I'm reading this correctly, so basically both obey the same rules, but you apply like, for example, this teal color or whatever, I don't know, like a light blue, right? So because it's at one, it just goes and it covers everything. So this one doesn't get applied. So maybe it's a thing where you sort of like, this would be this purple color every four, but then this is every single one. Okay. So I'm going to assume that's what it is. So we need a, so this is a 20 by 10. So we're going to need a 20 by 10 grid, 20 by 10. Oh, and just to be difficult, they threw in a yellow color here. So there's five spaces. Then we have our yellow color that doesn't go all the way to the bottom. So the gray one, this means every four. So it goes like this, then like that. The blue one means every two. So kind of like how they show it here, we go like that. And then the red one, every single one. That's great. So then the orange, the orange, since it's one, I mean, it would kind of be like this, right? So it's like that, but then we're going to cover the purple over it. So this one means it's every two. So kind of like that. And these are purple. And then we have this one was two purple and like that. So two means it goes every two. And this means it goes every three. So like every three like this, basically. So it goes one, skip two, one like that. So one, skip two, one, skip two. All right. This is it. Let's go. Oh yeah. Winner, winner, chicken dinner. All right. So they seem pretty straightforward. They're a little bit complicated at first, but if you kind of just sit there for a little bit, it kind of like you start noticing the patterns. And yeah, it's one of those things where humans would be pretty good at this. So it sounds like humans beat all of them, or at least capable of beating all of them, but large language models might struggle on figuring out how to do it. So I'm not going to go through all these, obviously, but just from the few that I tried, it seems pretty good. And as they've mentioned earlier, efficiency will also count here. So by intelligence, they're no longer just looking at, does the capability, the accuracy, they're also looking at efficiency. They can no longer report performance as a single metric. So this leaderboard will also now track the cost of the performance. So the reasoning models were sort of adaptive to the test because you can crank up how much you would think about any given problem and thereby improve the score. This new test is resistant to that sort of ability. By the way, let me know in the comments what you think about this. Is this, you know, quote unquote fair? Is this good to limit how much time you can spend thinking? Obviously, as we improve the efficiency of these models, it'll cost less, but is that a good metric? The other interesting thing is it's no longer just about climbing the leaderboard and trying to get the highest score possible. They've adjusted the scoring to incentivize conceptual breakthroughs. There's a $75,000 prize for the most significant conceptual contribution, a $50,000 prize for the highest score. The grand prize has been increased to $700,000. So definitely recommend that you go and click on play. It gives you a daily puzzle, which by the way, is timed and that does get reported along your score. I think you'll be happy to know that I'm still more generally intelligent than an AI and I have my certificate here to prove it. They cranked up the difficulty just a little bit. Again, not in terms of like you need some massive processing power. You just need to kind of pay attention and try to figure out what the patterns are, what the rules are. And here's our current leaderboard. So as you can see here, the human panel, well, they got 100%. So out of the people that are taking, I think each question, at least two people solve it. So I think they're using kind of a group of people and at least if you solve it, then it counts as being solved. So all the questions that are included were solved by at least two people in that group. Then we have the O3 low chain of thought and a synthesis, right? Comes in at 4% and costing $200 per task. And it kind of drops off from there 3% for the O1 high. Architects, which received the ARC price from 2024, gets a 2.5%. But that's the first model here that actually comes under the cost per task. DeepSeq R1, the Chinese open source model, noticeably is also here kind of high up there. 1.3% and just $0.08 per task. So definitely a good showing. There's some rumors about the new upcoming DeepSeq reasoning model doing very well on the RKGI. I can't, you know, verify those rumors. So they're just floating out there. I don't know what the source is. So maybe, maybe it's true. Maybe it's not. We'll see. We are still expecting DeepSeq. I'm guessing they're going to call it R2, the next iteration of this one, to be dropped pretty soon. So exciting to see a fresh unsaturated benchmark, as they put it. Again, the best AI is 4% and way above in the cost per task. The first opening AI model that will qualify is at only 1.7%, right? So we have a lot of room for them to improve, to go, to climb this leaderboard, to get to the 85% needed to qualify for the grand prize. Now, this is from the manifold market. So we're able to kind of place bets on various outcomes here. They're asking, will the ARC AGI grand prize, so that's the first version, 2024 data set, is anybody going to claim it by the end of this year, 2025? So the market thinks there's a 27% chance. It looks like there's another betting market here. If anyone will score 70 plus percent on the ARC AGI 2 within three months of its release, people are thinking there's an 8% chance. What do you think the chances are? What would you guess? But here's some food for thought. So this is Isaac Liao, machine learning PhD, previously computer science and physics at MIT, international physics Olympiad 2019 silver. Smart guy, I would say. I think that's safe to say. I'm going to give him a follow. But, so he reposted the ARC AGI prize too, but if you keep scrolling, he's got this. Introducing ARC AGI without pre-training. No pre-training, no data sets, just pure inference, time, gradient, descent on the target ARC AGI puzzle itself, solving 20% of the evaluation set. So interestingly, we have a lot of different approaches, people trying out new stuff. The organizers of the ARC prize themselves, they're not really doing this to, you know, have everybody lose. They are looking for somebody to actually come up with the correct solution. Francois Chalet is giving some sort of a hints about what he thinks might be the right approach. But of course, the whole point of this is for everybody all over the world, everybody that wants to compete to try this out, maybe come up with something new and innovative. It has to be open source. So whatever new things we discover from it, everybody else will be able to kind of learn about those and try them out. So let me know what you think about this whole thing. Do you agree with them limiting how much a compute sort of you can spend on it, what the budget is? This seemed kind of low, 42 cents per task. Do you agree that that should be a restriction? I mean, they still had their $10,000 sort of total compute that was previously. But let me know what you think about that new restriction. What do you think about the test in general? Certainly, they seem well organized to kind of push innovation forward for people to come up with new ideas, to be incentivized, to put them out there and to hopefully, you know, kind of an open source, a group effort to try to push this thing forward. I got so excited about being generally more intelligent than an LLM. But I think I jumped the gun there. Because at the end, I said, are you starter than an LLM? Which of course should say, are you smarter than an LLM? So I think I won the battle, but I lost the war. But check it out for yourself. Play the game. Tell me what you think about it. If you made this far, thank you so much for watching and listening, and I'll see you next time.