Open transcript
We'll have to haze him once he gets on here. Go live so we can talk shit live. It'll be perfect. All right. Looks like we are live. We're six seconds in. Let me see how many people are on. So, yeah, everybody pile it in here, please. You're fired, Joe. I had it all set up for you perfectly. Text messages, everything. We're live right now, so let me just make sure that everything's going well. It's all good. Joe got good Mexican food, though, yesterday. At Mexican food place, Adelita's on Kirtner, San Jose. Super good. Can you hear us, Joe? Mm-hmm. Okay, great. Very, very cool. All right, let me make sure audio levels are good. People in chat, please let us know how the sound levels are. There's a subsection of the audience that really loves to mess with, not just me, but anybody that's live streaming, by messing with them telling the sound is off. Oh, you get haze, too? I get that stuff also all the time. I can make people go deaf. They're like, I still can't hear it. Turn it up. Turn it up. All right, so audio input capture. Okay, so I, it sounds like you guys are a little bit more loud than I am, but that's your mind. Okay, so one said the middle guy is a little loud, and it's true. I'll go down slightly, so I'm not going from yellow. Oh, yeah, let me know. I can, it looks like I can also adjust it on my side. Yeah, so everybody, you know, thanks so much for joining us on the live stream. We still have people piling in, so let's give everybody a few minutes. But we're going to kind of start chatting, and in a few minutes do like a introduction for everybody. Let me just make sure that people are piling in here. Chat is slowly beginning to wake up. Nice. How's everyone's weekend? Chat? Did you guys have a good weekend? Did you watch any good movies or anything? Play any good video games? Anyone here play Manor Lords? I love Manor Lords. I went in for a new patch. Joe's like, I don't play those type of games. I play Dwarf Fortress, where you just constantly get your dick kicked in 24-7. For pure frustration. Exactly. Yeah, Dwarf Fortress is great. I tried it out. Very complex. They made RimWorld, which captured a lot of the feel of it, but just a little bit less complicated. I enjoyed RimWorld for a while. Do me this one thing, Wes. Put your mic closer to your mouth. Yeah, definitely. I hate the fact that sometimes you hear all the mouth noises. I don't know if people enjoy that or not. It just, I guess, depends. We have that smooth DJ voice. Hey, ladies, this is Wes Roth at 96.5 KOIT. Another light minute music set for the next 90 minutes. You know, hit that like and subscribe. 102.7 AGI. I know. All right, so let me see. Let me see if I can ask a trusted source. Do you, okay. You know what Microsoft likes to do is they like to update my mic settings behind the scenes. And so I always try to check those because it will go from 90 to like 60. And it's very frustrating. So you might want to just check one more time. Make sure your mic settings are good because I'm seeing some people who are complaining. But they could also be deaf. And we have nothing against deaf people. But, or no, people of hard of hearing. Let's see. Yeah, I'm at 86, so I should be good. 86. 86. Yeah. Okay, so everything should be good. I think now that we've adjusted it, everything should be in line. This is, my mic is turned up all the way. All right. Yeah, look at that. Chat is warming up. We've got tons of people. So, yeah. And we're still testing the sound. So we'll change it as we keep going. Just give us a few minutes here to double check everything. Someone just complimented my microphone. I feel so good about that because when we first started, we used to use the internal webcam on my microphone. And it was completely sketch. And you would see the back area of my shed. And for some odd reason, every time at like 11 o'clock at night when India would wake up, they'd be like, great episode. But the bald-headed brown guy, his shed is so messy. I thought he works at Google. I thought he was smart. He's just so messy and bad. I was like, Jesus. So I just turned around to this now. And then people now complain and say, I want to see your messy background. So you just can't make people happy. So let's get to the AI stuff, Wes. Get to the AI stuff. Oh, actually, should we introduce ourselves? Yes, please. Can you guys please introduce yourselves? Give us a little bit of a background. Meanwhile, I'll kind of do my little pre-flight checklist. Awesome. My name is Jordan. Me and Joe work in the Svick Podcast. Before the podcast, I worked at Google for 10 years. I worked with Joe there, too. I worked in M&A for eight years at Google. And then I left Google because of that 200,000 heads. I said, I'm going to go to something smaller. I just went to Slack, 2,500 heads to do M&A. And then Mark Benioff was like, no, you don't. And they acquired us. And so I went back into a large company and wanted to start crying. And after that, we started Svick Podcast. And then ChatGPT got released when I was at Salesforce. So I tried to integrate where I could. Joe, let's hear your background. Yeah, my background is a whole bunch of tech companies. I sort of started. If no one is loud enough, some people are saying, I mean, are you sure your sound is turned all the way up? Because I feel like we're all the way up on everything. But we'll keep going. My gold standard is if you just look at the Riverside screen, you'll see audio bars, audio voices. And I'm hitting yellow. So we should be good there. And that worst case scenario, we can also release this and do an audio touch up. Yeah, exactly. And mine looks like it's all the way up. So it looks like everybody's looking good. Anyways, okay. We're finally. Yeah, so I think we're pretty much ready to go. I think everybody's piling in here. So, yeah, thank you so much for being here. We got a couple of great things to talk about. I just realized I don't have my notes. So, but, yeah. All good. Yeah, yeah. Yeah, and as I was saying initially, the chance. There's so much tech stuff that we're using here from Riverside to OBS to a whole bunch of stuff. That the chances of this live stream going perfectly without any tech issues. There's 0% chance of that. I'm just warning people in advance. Right, and we're also all on ketamine, too. There we go. So, I mean, you know, just the way it is. We're high. I also created a presentation for us later on. Oh, cool. Using Gamma, going over scale AI, and then all this acquisition craziness. But we'll love to hear what you want to go through, Wes and Joe, first. Yeah, absolutely. Well, yeah, Joe sent some excellent papers. I only covered one of them before. So, we definitely want to kind of go through it, talk about those. How interesting is the stuff that scale AI is doing with Meta? Is that something interesting or not really? Mm-hmm. Okay, all right. And, yeah, it definitely seems like there's more and more of an overlap between kind of this is more like DOD and all the Tech Valley stuff and war and stuff. They just, yeah, definitely more and more overlap. So, maybe we want to touch on that because you are, you guys are very much into, you guys are kind of like the insider. So, for a lot of the people on the live stream, I think this is going to be a real insider look. And, again, obviously, your opinions are not that of Google or any companies that you've worked for. Obviously, this is your own stuff, right? But we'd love to hear your opinion on all of that. We're confirmed conspiracy theorists. Exactly. We have our – I just took off my Star Trek uniform and was going to talk about – Oh, you should have worn it for the show. No, no, just – I'm going to go old school like David Shapiro and talk about – Are you a security guy? Is yours gold or red or blue or what? Mine's red, dog. Red team, dog. Red team? I know some people are blue team. I just want to know who's going to make it through the episode. No, I'm red team. First to go. First to get killed. That's just the way it is, dog. All right. I'm not part of the Screen Actors Guild, unfortunately. I've got to get killed early. So, what you can do is I create a presentation real quick because I'm a corporate sellout. And let me get that thing up real quick. Okay, so – Meanwhile, while you do that, I'm going to mess with my mic, so kind of ignore me. Okay, cool. And then Joe, Wes, feel free to just like cut me off anytime you want to bring something in. Did the model pick this photo for you? Yes. Well, actually, take it back. I said Mark Zuckerberg, and it came up with this with his chain. Outstanding. So it's in their stock or something. That is really impressive. Yeah, exactly. So – and then I have a live stream comment saying this sick guy blocked me from his channel when I talked bad about Sam Altman. No, I probably blocked you because you're an asshole. Okay, so MetaX Scale AI, the $14 billion power play. Meta is finalizing a $14 billion cash equity deal for Scale AI. This is Facebook's second biggest M&A after WhatsApp. So when they bought WhatsApp, they spent 10% of their market cap for that company. Dude, what was your reaction to the WhatsApp acquisition at the time it happened? Well, because I was a stupid American who never touched the glory, which is WhatsApp. And I'm like, what the hell is – what are they doing with this deal? It's like a chat messenger. Why not using text message? And then I finally got tons of awesome Indian people moving into my house because I rented out like in Silicon Valley. And they were like, if you want to talk to me, you're going to use WhatsApp. And then now I use WhatsApp for everything, and it just completely makes sense. And then finally, Joe, the news came out that WhatsApp is going to be monetized now. They're going to finally start putting ads into it. Yeah, you've been advocating for that for a while. You've got to monetize that, baby. I mean, it's the way it is. And so – and I look at it as Zuck is basically saying, hey, you know what? If we're putting LMs in all this, this inference cost is going to stack up on me. And the pre-training cost and all the capex I spend, I've got to figure out some way it justifies to the stream. So anyways, the scale AI deal, people are like, oh, they're paying $14 billion for Alexander Wang. But if you look at it as a percentage of market cap, it's less than 1%. So this thing could – and yes, I know there's like a couple trillion worth now compared to where they were. But this thing could implode. It's still like a rounding error for Facebook. I think it's a good risk. High upside potential, minimal downside. So Alexander Wang is to lead the new super intelligence team that Mark Zuckerberg is recruiting. And then he's also calling engineers and offering them the eight to nine figure salaries just to like, please join me. And then also this is to basically get them back into the game after the cuck-tastrophe, which was Llama 4. Like Llama 4 was just, oh my god. I mean, Llama 2, I literally was like showing rap videos and saying that's Mark Zuckerberg because he has so much swag and he's so cool. Or as a kid says, so lit or something. I don't know. I'm an old millennial. And then Llama 4 came out. It's like what's going on internally? Like Joe, what happened? Why did that turn around so negatively and so quickly? It felt like they were trying to beat some internal – or not internal but public benchmarks. And they maybe stretched a little bit, cheated a little bit. I don't know what the right word is. And then they got a very negative reaction because they advertised that they did well on these benchmarks, but the model didn't seem to do well on anything else. And so I think people lost their enthusiasm. Right. And then shout out to machine learning street talk guy. He basically did a 25-minute video saying like this is all – this is Goddard's law or Campbell's law for social science nerds where you basically say, okay, we need to increase employment. Let's focus on that number. And then people will play games and say, okay, I'm going to hire people to like start digging ditches with spoons. And so it's all the same thing here with these benchmarks. But people are probably wondering like what's going on in M&A? It's all confusing. Jordan like shed some light on it. And I just got back from Holiday Inn Express, so I'm going to give my best take at this. You get three major M&A deals, Aquahire, license and release, and stock purchase. Now, Aquahire is basically a glorified hiring exercise. So anytime you hear any of your homies or whatnot try to like shine and say, hey, well, I got acquired by Google or Facebook, you should ask them was it an Aquahire deal or a full stock purchase deal. If they say Aquahire deal, that basically means they didn't get really paid anything. And Google said you have no enterprise value. We don't want your company. We just want to hire all your employees out, and then we might give you some additional retention money, which is golden handcuffs. I don't know if anyone here has watched Silicon Valley on HBO. Joe won't watch it because it's like the story of Dorian Gray where if he watches – He gives him PTSD. I watched one episode, and they did an M&A negotiation, and I was getting PTSD from it. But there is a scene where all the founders are on top of the roof, and they're just playing video games, and they're like, wait, what are you doing here at Hooli? They acquired you, and you're not doing anything. It's like, oh, we just rest and invest. We stay here for three years until our retention pays out, and then we leave. In the meantime, corporate doesn't want to do anything with us, and so that's where we are. And when I was an admin with Joe back in the day, one of our directors I was also an admin for, he got acquired, and for some odd reason, the Google product team didn't want to do anything with him, even though he was like a genius. And so he just sat in his office and vested, and he would go mentor kids and things like that, and he would raise his hand every time. He'd be like sometimes with the VP saying, I know what you're trying to do, but it probably won't work. And they would ignore him, and then the thing would fail anyways because he didn't listen to him. So anyways, most of the deals back in the 2010s were aqua hire deals because most startup ideas fail. And then you have these license and release deals where basically it says, okay, we want your team, and we also want some of your IP because your IP is decent, but we actually don't want your business because maybe you made hot dog not a hot dog app, which is just stupid. But it had a nice image recognition model under it that we're going to try to reapply to something else. So back in the day, those license and release deals, you maybe got some money back on your dollar invested. So a VC in an aqua hire deal gets nothing, but for license and release, they might get $0.85 a dollar. If you're lucky, maybe you might get maybe $0.10. Maybe you might get your full dollar back or something. But then things change in the regulatory environment, and we'll go into that later. And now we're seeing all these AI companies saying, hey, license and release is back in vogue. So we're going to start doing those deals to get away from regulatory heat, which I'll go into in the next slide. Then you have full stock purchase deal where everyone gets paid. When Salesforce acquired Slack, when Google buys Wiz, that basically means we're going to pay a gigantic premium on your equity. So an investor is going to get maybe 100x or 1,000x or something ridiculous. And those are the big, awesome deals when people get rich. So let's go to the next slide here real quick. And then Wes, if you have any questions or you want to stop me, just let me know because I'm one of those talking heads, majored in political science. I won't stop talking until you give me money. This is great. So just a really quick audio check. Is this better? So now I'm using this microphone. You're still saying, bro. Sorry to be that guy, and I hate doing it because I was that guy for the longest time. But we'll have to do something in post. And also here's another thing. Some people listen to their speakerphone, so it makes it worse if you're not listening to your speakerphone. It's better, but you're just a little faint. But feel free to see. Okay. I think if you two can hear me, it should be fine. I think I'm doing a different mic for the stream versus if you guys can hear me, this should be good. And then, Joe, I'm sorry if I did. Did I cut you off during your intro? I'm sorry if I did. No, no. Okay. I'm good. Okay. Sounds good. Yeah, everything. Gradient check says Joe could use some color in his background. Yeah. Why don't you have some color in your background? Some pink, some blue, you know, a nice painting. Actually, Joe has a Japanese watercolor that the Tokyo office gave him because he's so awesome. It shows Google, and then one O is a rising sun, and the other O is a Japanese version of Joe, watercolored. It's me, animated. Interesting. Before speaking to my admin, that was sent to me before I met him, and I saw this, and I was like, who the fuck is this guy? Who does this guy think he is, stealing the Google logo? This guy's an asshole, you know? And then I met him, and I'm like, ah, goddamn, now I'm going to make a watercolor for him. Okay. Okay. So why licensing deals fly under FTC radar? A lot of things for FTCs, it's based on the antitrust laws that were used to break up the railroads because back in time, the railroads, oil companies from the Gilded Age, people thought they had too much power, and they're using that influence to squeeze small businesses and whatnot. And so what the FTC looks at is if there is a market of, let's say, search, for instance, we all know that. If Google went to go acquire another search player like Bing or something, which wouldn't make any sense, the FTC would say, hold on, Google, you already have like 80%, 90% of the market. You then bought Bing, and let's say you got a few percentage points more. The FTC would say, oh, well, you kind of already do have a monopoly in working on that, but now this is like a super monopoly, so we're going to prevent that from happening because the theory is you're going to raise prices or something. But when you do a license and release deal, you're not buying the company. You're just taking maybe some of the IP and the headcount, but that organization still exists, and it can go off and die for all you care, but what's important for you is you're not going to get the FTC review like you used to do. Now, the FTC could change and say, well, in the Scale AI deal, yes, Meta owns 49%, doesn't have a majority, but for all intents and purposes, Alexander Wang is still on the board. You're probably going to get board seats. Alexander Wang is going to get even more money to stay at Facebook, probably in stock. So do you think he's going to be aligned with what Facebook wants for the future of Scale AI's company? Yes. So they could say it's a de facto acquisition, even though it's not 49% and they could investigate. But we'll see. Then when you do these type of deals, you get regulatory fast track. So I don't know if you all saw three days ago, Google announced they bought Wiz months ago for $32 billion, which was that deal happened. They started that company like three years ago, and then they get a $32 billion payout, and that was their second acquisition because they got acquired by Microsoft back in the day. So now their only issue they're dealing with now is they want to get individual, not just yachts, they want to get yacht craft carriers so they can have supporting yachts and helicopters coming in. But the issue with that deal is news just came in that – I should not read the livestream comments while I'm talking. News just came in that FTC is doing a review on Wiz. So now they have to go through a year-long approval. And the FTC does not like tech right now. The Republicans hate them. Democrats hate them. And just like the Figma Adobe deal, you could see things fall apart. Now, the Adobe deal fell apart because the UK's CMA, which is their version of the FTC, was going to block that deal. And then so Adobe said, you know, we're just going to walk away and pay a gigantic break fee. I think it was maybe the hundreds of millions but not billions of dollars. No, it was a billion dollars. Billion dollars just to say, you know what? We actually paid too much for this deal. Here's your billion dollars. And you probably – They actually did wildly overpay, and their stockholders were very happy when the deal fell apart. Joe, go into that because you worked at Adobe back in the day, and you know the industry. Well, it's amazing that Adobe wasn't able to compete with Figma. I mean, they had a huge amount of warning. You know, the apps were moving online. They were becoming collaborative. You saw it in all the productivity apps. And Figma's first product was a drawing product, which really would have competed with Illustrator on the Adobe side. And it didn't really take off in the way that the Figma founders wanted. You know, they wanted a much bigger business. So they went back to the drawing board and came up with a more design-focused product that became the Figma that we see today. So Adobe had at least two years of warning that this thing was coming. They spent a bunch of money and built their own team and probably over five years tried to build a competing product. I forget the name of it. It was something MX. But it never managed to generate much enthusiasm. They eventually shut it down when they did the Figma acquisition. And then when they acquired Figma, they paid a tremendous amount of money. I mean, Adobe is significantly smaller than Facebook, so I'm guessing it was several percent of their market cap. And I think the amount they paid was just below the threshold where the board of directors actually had to vote for it. So, like, the CEO and his team sort of put together this deal that would just get under the limit and they didn't need approval because I'm pretty sure they would not have gotten that approval. And then when the deal fell apart, everyone kind of rejoiced and their stock recovered. Right. No, well said. Yeah, they definitely were overpaying that one. And a good sign of if you're overpaying for a deal is how badly Wall Street will hammer the acquiring company stock. You'll see it usually get crushed if it's way too much. Yeah, so just as an exclamation point on it, Meta's stock is up the last two days. It's up significantly. Like, people are rejoicing that Zuckerberg is doubling down on this AI goal. Right. Two good pieces of news. One, he's restructuring his AI org. And two, he's monetizing WhatsApp. I know everyone here is like, oh, you corporate sellouts and blah, blah, blah. We're just telling you what Wall Street thinks about what's going on here. Yeah. All we care about is if you like and subscribe to Wes's channel, if you want to, our channel too. So second to the most part. You sell out. I am a corporate sellout. Look at me and my meth lab over here. So the other thing about license and release people I'm talking about is no liability from acquired company. When you acquire a company, you acquire its whole entire legal history and its liability. And from – they could have done some effed up things before you acquired it. You still are going to be on the hook. And now lawyers are going to say, well, it was a company worth $100 million. Yeah. Probably don't want to go after it. Well, now it's owned by Google. It's multi-trillion. Hmm. This could be interesting for a class action lawsuit. So as a side note, one acquisition I worked at at Google, it was one of our quickest closed deals. We got notified of the deal, and we closed it within like two to three weeks, and it was an acqui-hire. And it was from a company that's called Homejoy. And basically what they were doing is they were going to be the Uber, but for cleaning services at your house. And this is before Uber went to war, the state of California, and got the proposition to classify Uber drivers as contractors and not full-time employees. Homejoy wasn't there yet, and so they were doing well in their business but did not have as big pockets as Uber did. And so they were right in the middle of raising their next round of funding, and then the state of California raised a lawsuit against them. And so the woman who was organizing it, I forgot her name, but she was awesome. She then effectively was like, okay, well, I guess we're not going to have any funding because all of our VCs pulled out. So then she came to us, and we were able to do an acqui-hire where we were like, hey, we want nothing to do with the state of California lawsuits, but we will make sure all of your employees land here at Google and get good jobs. And so we closed that one pretty quick, and it was one of my favorite deals because we gave all those people good jobs, and we didn't give them pink slips, and they were able to pay their mortgages and do great things at Google. So let's go to the next slide. Now, license-release deals are cool again, like I mentioned. And so here are the major ones. We have Microsoft Inflection license-release deal. That was a $650 L&R. And that was one of them. And then we have Google and Character AI. That was for $2.7 billion. That's where Google was at, Gnome Shurier, and a few others. Gnome was one of the authors on Attention is All You Need, and he's a really, really good engineer. Got about 30 researchers from that. And then Amazon did the Adept deal for $330 million. And then we had the Metascale AI deal, which is an investment, but it feels like a license-release because they're also getting the CEO plus a handful of employees. So let's go to... This Metascale AI deal feels a lot like the Google Character AI deal. Yep. Why don't you go into that for a second? I mean, as you said, they're sort of extracting the CEO, founder, and a couple of key researchers. That part's very similar. They're structuring it in a way that sort of avoids regulators. That feels the same. And it's an incredible amount of money for what looks like an acquihire after you sort of look at all the components of the deal. Exactly. But now we'll definitely have to go into the valuation point you mentioned. I have a slide on that, but you're right. It looks like it is an incredible amount of money. So Scale AI, founded in 2016 by Alexander Wang and Lucy Goh. They do data labeling and evaluation, and they have two different suborganizations. Is Lucy Goh staying with Scale, or is she coming to Meta? It looked like she was staying with Scale. Interesting. Yeah. Yeah. I didn't – because she's pretty prominent, so they would have mentioned – oh, yeah. And also my co-founder, Lucy, is joining us because he sent out a note, and he just mentioned, hey, I'm still going to be on the board of directors of Scale. And very little is going to change. We're going to get our chief of staff to lead – he's going to be the CEO. He's the chief of staff or strategy person that's going to be the CEO. But nothing about Lucy. So they have two suborgs in Scale AI. One org is basically – they do PhD-level data training sets that these LLM providers pay for, like OpenAI and Google. Well, it used to be Google. Google's now signaling because of the deal. They're backing out. And then they have another – And so is everybody else too, right? Everybody's bailing out from using them, right? I – this is breaking. Did you hear any other companies? I think so, yeah. I forget who, but like all the big players, I think, signal that they're getting out. They're not using Scale AI. Yeah, OpenAI, from what I heard, OpenAI says we're cool with it. We're going to stay. Okay. It'll be interesting to hear what Amazon's doing or become a major player. They're – they've been kind of sleepy over there, but, you know. And then you have the other business, which is more of like Amazon Mechanical Turk, where it's like 100,000 employees, and they can help you – contractors and whatnot. They can help you with data labeling for just other types of data sets and whatnot. And then so let's go to the next slide. So what is Meta really buying? Just look at this as getting Alexander and energy into the organization so they can hopefully turn things around because it just became kind of a clusterfuck during a lot before. And then they're going to have the ability of using Scale AI's data pipeline to help them train models. But what's interesting thing is me and Joe have been talking on the show for the last year and a half about synthetic data and saying, okay, there's more and more synthetic data coming out. There is now RL with verified rewards. There's now models coming out that you can give it a confidence level. They can say, you know, I'm going to trust my gut, and they can actually increase its performance to a degree. Like this has to have downstream effects on Scale AI's business. And so we saw some of that in regards to Scale AI missed its target, its revenue target last year. It was supposed to hit a billion dollars instead of hit 870 million. I mean, to us mortals, I'm going to get still a lot of money. But it wasn't the growth as fast as they expected. And so they were getting some pressure from that. And then also if you look at the valuation, the amount of money that Facebook put into this deal compared to their previous valuations, it wasn't such a big of an uplift in exit value, which made me think of, oh, maybe they were realizing we might get to a point where we're getting to the end of the line of this type of business. Joe, you made some good points during our conversation about enterprise sales cycle. And maybe they were seeing something from that, too. Maybe you can go into that. Yeah, I think there's, you know, a lot of ways to signal that deals are delayed. And so if people start to think, okay, maybe I don't need this data set, or maybe there's other ways to generate this data, or maybe synthetic data is going to become available, or I can produce my own synthetic data, or any of those things are true. A typical thing in enterprise sales is that you'll just see the sales cycle get stretched out. So signing the deal will somehow just get delayed. And then salespeople are looking at their set of ongoing deals. And if they see that they're all tending to get delayed longer and longer, that's a really bad sign, right? And if your book is getting delayed that way, you kind of assume that things are slowing down. Many of those deals won't actually go through, and that your sales cycle is somehow being impacted by something in the environment. And that's a sign that either the product's not right, or the economic environment's slowing down, or something's going on that's big. And if that was happening to sales, like at Scale AI, I could easily imagine the founders getting nervous. Exactly. And so you're probably thinking, okay, maybe it's time now to get out while the getting's good. You know, we can't take this anywhere higher. Right. Maybe this is the peak, and it's the right time to do a deal. Exactly. Gotcha. I look at this as a mega, ultra, super success for Alexander, getting to this point. I see a lot of hate on Twitter saying, you know, this company has no value, and all they do is contractors or something. I'm like, oh, okay. Go get yourself a $15, $14 billion exit. Oh, you can't? Then go back to Wendy's. So I think everyone should, you know, be applauding what happened here and not stressing Facebook so much on, oh, they spent $14 billion on him. When it's like, hey, they have this capital. They need to reinvigorate their AI org. If they're able to get that show on the road and get things going, then that's a huge opportunity for them to increase revenue if it's done right. So let me go to one more slide here. So kind of like risks on the horizon is just Google walking away from a scale AI contract, which I don't think Facebook really cares about because they're more focusing on Alexander and crew. But other customers could leave. And then there's a question of are they overpaying for the price? The multiple is about 12, 15X on their revenue. We don't know if the rest of this year if the revenue could hit a billion or be less than that. And then the other thing is just cultural mismatch. Alexander is 28, works at a small company, knows how to move things and ship product and get stuff done. He then goes to Facebook, which is a gigantic government and works through politics. It's very hard to get stuff done. He might have an alignment with Mark, but then getting what he needs done through the rest of the organization could be like pushing a safe through sand. So you have to be seen how that's going to work out. Joe used to work at Facebook. Any thoughts on potential stumbling blocks Alex might deal with going into that organization? Well, I think there's always risks in a big organization like that. You have people who are already there and sort of consider that this is part of their charter. I mean, Facebook already has a couple of teams working on AI and ML. There was just a release out of Jan LeCun's team just, I think, the day before yesterday. So that's probably the biggest issue. And then, you know, from his perspective, is there going to be internal movement? Like are people from other teams going to join his team or are people going to try and recruit away his best people? So I'm sure there's jockeying for position going on. And then lastly, Facebook itself is still under investigation by the regulators. Right? So it's a kind of uncertain environment for them. Right. Because everyone here has seen the movie Office Space or Half-Baked by any chance. Mm-hmm. And there's that scene where the guy is sitting down in prison and then Squirrel Master comes up and goes, Hey, he's my bitch. Like, don't mess with him because all of the prisoners are trying to attack him. Yeah. Joe, you had – yeah, I don't know if you can – well, I'm putting you in the spot. You had a similar experience in your onboarding at Facebook where you're getting, like, recruited. Maybe you can tell a little bit about how they have this weird way of allocating engineers to certain teams at Facebook. Yeah, I think that's actually one of their strengths. They have a really interesting training program they call the Boot Camp. And so new hires coming into Facebook are not normally allocated to any specific job or team. They just come in. They go to this Boot Camp training. They are there for anywhere from a couple weeks to a couple months. And they go through training sessions where they learn how to do things inside the company, including checking in changes to some of the products, like, very quickly. So usually they have a goal of, like, first day you check in a change to one of the products. And then on the flip side, all of the teams inside Facebook package up small changes or bug fixes that they want, and they sort of curated them in a way that new hires coming in can pull something off the queue and then go through it as an initial project. And then there's someone in that team who helped file that bug or that issue is listed on it. And if the person who's coming in as a new hire and working on that bug has questions, they can contact that person. Well, this is a perfect moment for the existing employee to sort of see if the new hire is someone they want on their team. Right? And they're doing it in an environment where it's, you know, low stress. Well, at least low stress for the existing person. And a chance to evaluate this new person. How good are they? How fast are they? And then if they see a person they like and that person is showing a certain facility, they'll try and recruit them on the spot. Right? And ideally for Facebook, the new people they're bringing in are getting recruited by two or three internal teams in their first couple days, right? Because that means they hired a great person, and it means they're showing a lot of skill, and it means the teams have a constant stream of new hires coming in. But it's weird for most people because they've separated the hiring and recruiting from the placement, the assignment to the teams. Now, tell the funny story about how they were fighting over you at lunch. Well, when I went there, I already knew that I was going to build a team to work on privacy. And so I was, you know, in a so-called allocated person. I already knew what team I was joining. But I still wanted to go through boot camp because I thought it was an interesting possibility to, like, learn all their systems and also to see how their recruiting program worked. So I told them when I joined, you know, put me in the boot camp like normal. Even though I was going in as a manager and not an IC. And as I was doing, you know, bug fixes one by one trying to hit all the different systems that I was curious about, people would say, oh, you know about this stuff. Like, a couple of the bugs were around doing JavaScript on the client side. So I knew about that. And another set of bugs were around localization. And I had a long history with that from much earlier. And then a third set of bugs were around advertising, which I knew a little bit about from Google days. And so each time I would be fixing a bug, I would talk to someone who was listed. And I didn't really understand their system yet. And I would engage with them and ask all these questions about, you know, what's the issue? What do you think of these possible fixes? How should I pursue it? And inevitably, they would say, oh, you seem like you're interested in this. Would you like to join the team that works in this full time? And then I would say, oh, I already know which team I'm joining. And then they would be sort of frustrated, like, well, then why are you still fixing bugs and still in boot camp? So I was kind of, you know, going upstream at that point. There's one part where I guess he was sitting down to eat lunch and two managers kind of pounced on like, hey, so it looked like you need a team. And then someone super senior, like Skrullmaster came through and was like, nope, he's with me. Let's go, Joe. And both the managers kind of like screwed away. I think that's a core competence at Facebook if you're a team leader that you are aggressive about hiring because they are in fierce competition. And if you're really good at it, you are there in person. You come to the area where people who are doing the boot camp are usually sitting and you sort of mix with the teams and you try to put faces to the names that you know from interacting in the database. Yeah, that's an interesting thing. Yeah. How that's set up. That sounds like it would work very, very well. And I like the idea that you don't get just somebody dumped in your team. You kind of have a chance to kind of see what they're like working and then request that they get pulled to your team. It's kind of like almost like a second round of, I mean, it's placement, but it's a second round of recruitment almost like you make it to the inside and then you're recruited onto the team. That's an interesting approach. Is that common? No, it's very uncommon. And, you know, it's a fairly expensive thing for Facebook to do, right? Because it means all these people are delayed from officially having their job and their assigned tasks for, you know, up to a month while they're in this boot camp. And, you know, they pay all these people to be trainers. So there's like people walking around and then people from the various teams give classes, like hour-long sessions on how the different systems at Facebook work and how to do various kinds of work at the company. So that's also a sort of expense for them. And then there's this huge area where all the new people are attending classes and, you know, going through their projects and getting help from sort of hall monitors and so on. And that's also expensive. And then all the teams are carefully curating bug reports and improvement requests and putting them in the database, knowing that they, you know, it's not like they're setting aside work for themselves. They're purposefully setting aside chunks that can be done by a new person, right? So they have to carefully make sure all the information is in there and it's complete for a person with no context. All of that is very expensive. And I credit Facebook for putting in the investment. A new person coming in has a tremendous opportunity to learn all about the systems, all about the teams, how the company operates, how things get done, literally make changes to a running product on your first day, which is extremely dangerous. They're very proud of the fact that, you know, new people coming in, bring down Facebook every couple months. Right. By accidentally checking in some bug. And then lastly, it's a great way for all the teams to get direct exposure to these new people and vice versa. And you eventually find their match. The try before you buy. It's great. And then also speaking about introducing a bug, I was seeing some podcasts and the director, she was given an assignment and she actually brought down part of Facebook's site to a certain geography because she did somehow denial of service attack on Facebook. And she was petrified. And the engineers came up and were like, no, no, no. You found a vulnerability here that we didn't, you know, think of. Now we're going to fix it. And now to compare this, Wes, to how Google does things, I used to go to Google's new or new orientation all the time because when we'd acquire a company, I was the head HR contact. And my job was to prevent the rest of HR from hurting my people. It sounds like these are mine. And so what you do for Google's orientation is you would sit like a week, week and a half long in just meetings and videos about like our orientation and what we're doing and things like that. And you'd very gradually maybe get exposed to the code base. Whereas day one at Facebook, like Joe was like in there like doing things. And a lot of people appreciate that, having the ability to just jump into the code stack and get to work. So, yes. What else shall we talk about? Well, so I do want to touch a little bit on all the incredible research that's been going on. So, I mean, briefly, I feel like this is we don't want to beat a dead horse, but how about the Apple paper? Can we briefly just talk about that? What do you guys think about that? You want to start on the controversial stuff? Yeah. Yeah. Joe, go first because I gave a whole class the first 40 minutes. Your turn. Your turn. I'm not sure what's motivating Apple. I mean, it's strange because everyone, well, investors perceive that they are going slow on AI. And it's not clear that that's because they don't think AI is ready for their products or they don't have their own story together internally or what exactly the problem is. But they're definitely lagging other large tech companies. And this is at least the second paper that I've seen where they mostly spend their energy pointing out how these AI systems are not good enough, not ready for prime time. And then people sort of react strongly at first, like, oh, my God, these AI systems can't do something important. And then gradually sort of realize, hmm, you can pretty much do these things if you just use the model a little better. Yeah. And then the paper kind of disappears into irrelevance. And I feel like that's happened at least these two times. So I'm not really clear on what Apple's thinking with this kind of research. Do you agree with this, Roth? Is this sort of the direction you perceive? Yeah. I mean, yeah, I don't understand what the point behind it is. I've seen a few papers like this or some of the blog posts and more papers. Like one was literally called LMs Can't Reason. And, yeah, there's so many things where I don't even know where to start because, number one, you know, if we're talking about if something can do something or not, like can humans run a four-minute mile, right? Now, if we see an example of a thousand humans failing to run a four-minute mile, does that prove anything? No. I mean, you can strongly suggest that maybe it's impossible. But you need one example of a person running a four-minute mile to say, okay, so that's disproven. And then so, one, you can show me a million ways LMs fail at doing something. Does that prove that they can't reason or can't think or whatever? That's one. Number two is what do we even mean by think or reason or anything like that when it comes to LMs? Because that's kind of a very human-centric thing. So it's like if you ask these people that say LMs can't reason, okay, what's an example that would disprove that hypothesis? Give me an example. Like if I get them to do something, what would prove to you that they can reason? If you can't come up with an example, then this conversation doesn't make sense. Also, in the paper, they have the river crossing problem, which is impossible for after, you know, N5 plus or whatever. Like basically at that point, it's impossible to solve. So the model probably says it's impossible to solve and the market is zero. So it's just there's so many things there. Also, why is, you know, being smart enough to realize if a model is smart enough to realize I can't do this problem within my context window, but I can create a tool that solves the problem, then builds that tool using Python code or whatever, and then solves the problem. Like, how's that not reasoning? Like, why are we? It's strange. Like why we choose that as the definition. And the previous LMs can't reason paper from a couple of years ago was not of Apple or somebody else. But usually what they try to do is just hit it at some limitation that the model has. So before a lot of the puzzles would be like, you know, it can't count the number of words in the sentence that it's about to say. Right. Because it didn't have reasoning yet. So it couldn't think through it, get that data and then count it. It had to, which humans can't either. Like, I can't predict the number of words that my next sentence will have before I say it unless I write it, you know, I say it first, count it. Right. So a lot of the problems were like, yeah, like furniture placement where you had to meet certain constraints. Now that we have reasoning models, of course, it'll 100% succeed at all of those. Right. So now they're strictly and the other paper also had, you know, the context window limitations. So it's just hammering it in places where the context window would fail. And this Apple paper is now 100% context window. So, yeah, it's just one of those things. It's like so many faults with it. And like, number one, what is reasoning? Number two, what would be an example of LMs doing something that would qualify as reasoning? At which point you would say, yes, they can do it. And then, yeah, three, don't just hammer their existing limitations. That seems weird. Yeah. I'd be much more excited if someone, like, extended the capabilities. If they said, here's a limitation we bumped into and here's the things we did to try and overcome it. And maybe one of them succeeded. That would make me much more excited. Like, that's a contribution. I like your question, you know, what would have convinced you that the models are able to reason? That's a great question. And then lastly, I would say to them, when you publish a paper like this where you say, generically, models can't do X, there's a real danger that the reaction is, no, you can't get the model to do X. And someone turned around immediately and got one of the models, I think it was 03 Pro, to do the Towers of Hanoi with 10 disks, which was one of the more difficult problems in the paper. And it did it correctly without tool usage, which is kind of amazing because that's a long sequence of moves. And then someone else did it with tool usage, I think, on an earlier version of Claude. So, like, one of the major examples in the paper was already disproven by some random person in the community within a couple weeks of the paper being published. Like, that's kind of sad. It means you didn't really put out much effort to prove your core thesis. Yeah. Yeah. So, yeah, there's a lot of problems. Yeah, I agree. Plus what you both said, old Google Plus parlance. I'm sorry. It's hard. Old habits die hard. When I see all of these different research papers attacking these NLMs or people going on the press tour like Jan LeCun and attacking them and things like that, it reminds me of this quote from Max Planck. He says, a new scientific truth is not triumphed by convincing its opponents and making them see the light, but rather because its opponents eventually die and a new generation grows up that's familiar with it. That's good. And I look at my niece right now, and she gets to play with ChatGPT and talk to it, and she's happier than, oh, get up. And I wonder what their generation is going to think 20, 30 years from now if they'll be so focused on these questions like that. And that's why for myself, benchmark and benchmark porn, I mean, in the initial days when GPT-4 came out and whatnot, seeing the benchmark leaps and how things improved, I was like, this is cool. This makes sense. But then as time went on, the games other companies were playing by doing like, okay, we're going to compare one shot with GPT-4 versus 10,000 shot with our model. Look how we've improved on this benchmark. Wow. And so I started focusing more on things like SweetLancer or other measures of, you know, Fiverr, for instance. When ChatGPT got released, their job postings went down 17% because people were saying, hey, I can use this model now. They don't need clip art or copy editors. Or Stack Overflow is seeing their traffic completely implode. For me, the big indicator is going to be more benchmarking SweetLancer where it's people are exchanging money not knowing they're exchanging money to an LLM because they think they're working with a real developer for a fixed problem. Now, to be clear, developers, I think there is a lot of job security for you all in the future. Engineering work is so much more than just coding. It's just framing up the problem, dealing with internal or technical complexities and things like that. I'm not one of those who threw all your jobs going away. I like those type of indicators more. The big indicator for me is when am I going to hear my friends who are working at all these different tech companies say, hey, we decided to forego headcount because I want an AI agent instead to take on a specific role. I have not heard that yet. And that is something big for me. So when I hear all these startups saying, oh, we're agentic this and that and this, it's like you're using an LLM to augment people, which is cool. But you're not replacing headcount with that. So anyways, that's my kind of tangent. Last thing to bring it all home. Do any of you remember what happened to Apple in 2012 by any chance? What the controversy was then? Okay. So Apple was like, you know what? F Google and Google Maps. F them. We're going to do our own. Oh, Apple. Right. And so I created Apple Maps. Apple Maps. And Apple Maps was like leading people into wrong directions. I organized an offsite for Google. We were going to go into the Santa Cruz Mountains. And I went to BevMo and got two shopping carts full of like high quality booze. Then I was going to have Armadillo Willys show up with a van full of ribs, brisket and everything. And then we had a park team that was going to show up and do geocaching. And you're in Santa Cruz and the summer is beautiful. Our driver was taking us there on the bus. And I was just looking at the route he was going. And I was like, I think we doubled back somewhere. And I went up to him and I was like, hey, you okay? And he's like, oh, I'm sorry. This Apple Maps thing is putting me in the wrong direction. And everyone in the seats, they started laughing. No, no, no. It's like download Google Maps. It's going to be okay or get to an Android. And that controversy was so big that the head, one of the senior VPs at Apple had to resign. And Tim Cook had a huge just sore spot for Matt because he's the operations guy. He wants everything to be perfect, fit and finish. And he probably never heard, he probably for that whole entire year or two years was constantly hearing people bitch at him how they got lost from Maps. And I think scarred into his brain was never effing again. If we're going to launch something on the iPhone, it's got to be perfect. And I think for the LLM side, they weren't able to get it to that fit and finish level. And they said, F it. That's why Siri intelligence is not going to be upgraded until like mid-2026. Which is kind of a huge, like, I don't even know if they're going to hit that number. So instead, if you can't beat the other companies, what you do is criticize them and start releasing your AI research papers and saying, oh, this technology sucks. That's where Apple is right now. They can't do, so they're criticizing? Yeah, that's my read. But I'd love to hear, Joe, Wes, what your thoughts are. Well, it's all yours, Wes. I mean, again, I have no idea. That suddenly seems reasonable. I'm realizing this is really good. We're running a little bit late. Let's hit a few of the more fascinating papers. Yeah. Because the intuition paper, Intuator, right? The learning to reason without external rewards. Let me click over here really fast. Well, actually, you know what I can do is I can just bring it here into the window. So I'll block kind of our beautiful faces or somebody in chat put our big, beautiful, smart. What do they say? Big, beautiful, bald faces or heads or something like that. Whatever that was. Yeah. Thank you. They call me Megamind. Megamind. Yes. Yeah, that's good. So learning to reason without external rewards out of Berkeley. And this is weird. So I plan to do a video on it. So for people that haven't been, that haven't read this yet, it's strange. It's weird. I don't 100% understand why this would be. Like intuitively, it doesn't really make much sense because it seems like instead of using some sort of verifiable outside external rewards, they try to look at how confident the model is and its abilities to answer a question. And of course, the more confident it is, that correlates to it getting the right answer more. Or if it's not confident, then it's... And confidence, I mean, they describe the kind of the math behind what they mean by confidence. It's basically, it seems like how many different sort of branching ideas it might have about how to answer it, where it's a little bit more narrow, that suggests confidence. If it's like, it could be a million different things, then it's not confident. But I guess they asked, what if we train it and the reward, the reinforcement learning reward, was it getting more confident on the answer? And somehow that improved its accuracy. So does that make any intuitive sense at all? What do you guys think about that? Well, you're using the internal confidence as a reinforcement learning kind of scoring mechanism, right? So the confidence doesn't really tell you how to train the model. It just says, this response was more likely correct. And then you can choose from the correct answers and the original questions to do RL training in another round of the model, right? So there's sort of two steps happening there. I agree with you, though. It does seem like you're getting something from nothing to use confidence to decide when the answers are more likely to be correct. You're like, you know, shouldn't I just know if the answer is correct or not? But before we saw this confidence-based mechanism, we saw things like self-consistency, where you would just sample the model, I don't know, 16 or 32 or 64 times, and then take the answer that was the most common, right? Which is also kind of strange. It's like, well, first of all, how come the answers are different if you just ask the same question over and over again? That's down to the sort of statistical nature of the model. And then second, okay, if there's a statistical nature, it's tending towards giving me the right answer more often, and the right answers tend to cluster together, whereas wrong answers tend to be more spread out, like they're wrong in different ways. That's kind of weird. Okay. And then I can use self-consistency to sort of pick out the answer that's more likely to be correct. Anyways, however I do it, self-consistency or this internal confidence, the end result is that I get the correct answers, and then I take the correct answers plus the original questions, and I do a round of RL. The model that I get out of that is stronger than the one I started with, which is also a little bit iffy. And then I can repeat that process because now I have a stronger model, and it's even more confident about even more correct answers. Right. How much of this, Joe, do you think is like pruning in a way of we do pre-prune? I think we're doing a lot of these models, and all of these techniques, and the RL side is like a nice gardener of trying to cut away the noise and helping the model get to the point where we can get the distilled information that we need so it improves performance. And let me know how terrible that analogy was. No, it's a great question. And I think pruning is a good analogy because you could look at what they call the pre-training, the sort of ordinary training on a large corpus of text data. And you could say that is giving the model lots of information about the world and many ideas, good or bad ideas. Who knows? Like it's internet data, so all kinds of different ideas. And the model is looking for patterns in that data. And we have to assume that correct answers are guided by that pattern formation. Right. The patterns tend towards correctness because that's also, again, related to consistency. And so what's in the model at the end of that pre-training is a bunch of good ideas. But the model is still stochastic, and there's also bad ideas in that collection. And so your analogy of pruning is like identifying the correct reasoning traces or the correct directions and emphasizing them, which is what reinforcement learning is all about, right? The reinforcement part. You're not eliminating the bad answers. You're just adding weight to the correct answers and letting the bad answers kind of just fade. What kind of a body blow would that be if somehow they were able to figure out down the road, okay, we've located certain areas on the internet where people's ideas are just so bad, it makes our models dumb? I think that's a huge amount of data curation going on. Some of the papers in our recent list were about just cleaning up the data set. And one of the ways they clean up the data set is by asking another model or, you know, ideally a stronger model to sort of look at the data and trying to eliminate the lower quality information from the data set. And that also improves the training quite a bit. They were doing that with Quinn, weren't they? Quinn did a really nice job of that. They had a very big pipeline. And there was a recent paper, I want to say, Open Thoughts, where they created a fairly large open source kind of data set using the same approach. A lot of filtering, a lot of quality checks. Nice. Yeah. And one of the papers we looked at, it's also training, changing the weights of the models based on kind of the synthetic data that it generates. We should probably look at that next. But when we were talking about like pruning with RL, I mean, that's kind of what, so Dwarke Ash Patel had several anthropic researchers on his channel where they kind of talked about this a little bit. Like one of the researchers said, okay, like what if all of the whatever abilities are locked in the models? Like once it's trained, it's in there somewhere in the latent space. And RL, I mean, instead of saying pruning, he said it kind of like lifts out the needed stuff. It's kind of the same analogy, but that's kind of what they were talking about. So it sounds like, yeah, there's a lot more there that then maybe meets the high at first, but with reinforcement learning, we're kind of letting it emerge the proper things. Because, yeah, it sounds like it has the data for or some signal for what's wrong, what's right. It just, and the reinforcement learning kind of like helps it strengthen that signal, which is, it's just interesting because it intuitively seems like it doesn't make sense. But obviously, this would reinforce the idea that it's already somehow baked in there, and we just need to kind of like get it out. It has a feeling of perpetual motion. Like it doesn't seem like you should get something with this approach. And then going back to the scale AI conversation, I mean, their business, as far as I understand, is producing large, high-quality data sets. And a lot of those data sets are used for this kind of reinforcement learning training. And this list of papers that we've been recently discussing is all about either creating data sets from scratch or understanding how to reinforce the model without a bunch of hand-curated data along the lines of synthetic data or internal consistency or confidence or whichever technique you like. But that sort of implies that there's less value in hand-curating very large, high-quality data sets, which is the scale AI core business. Right. Yeah. And so another paper, let me put it on screen here really fast. This is the self-adapting models out of MIT. And here it's really interesting because they're showing that we're able to… Let me, I'm not, maybe I'm not seeing, you see it, Joe? Am I crazy? I'm not seeing it. Oh, you know what? If you guys are looking… Is there a little box showing up on Chrome? Are you using Chrome or using Internet Explorer? Oh, are you looking through Riverside? No, it's going to show up on… Sorry, I'm an idiot. Oh, no worries. No, no, no. And there's going to be a delay. So you might not see it for 10 seconds because YouTube runs a little bit of a delay. But yeah, so people should see it on… Yeah, so I'm looking at the live stream. It should be visible. So yeah, I'm showing that. Let me go back to the top. So yeah, self-adapting language models out of MIT. Perfect. And so what they're showing is that these, you know, they have a great analogy here. So it's like a student that goes to school, reads all the textbooks, read, you know, all the lectures, and then writes their notes, writes down the notes, kind of compresses all that information to the notes, and then kind of using those notes, studies off of those notes, and that's very, very effective. And so what's interesting is they kind of are doing that, but the models are also able to, through supervised fine-tuning, do these self-edits and change, yeah, change their weights. You know, so basically changing kind of like the weights, like how the brain of the model works, so to speak, to be fine-tuned to do some specific task, doing it by itself, which is, seems really interesting because, you know, one of the things that we kind of talk about is that maybe AI agents, you know, the autonomous AI agents aren't quite just around the corner yet because these things tend to fall apart over long-horizon tasks. They don't have that long-term consistency, coherence. And part of that is, I think, one of the reasons is the models are static, right? So they're, you know, they're sort of, they're trained, and then it's like, you know, going to work, and then on day one, you have a certain, you have a certain brain, right? And it doesn't change, you know, on day 100. You don't update your information, knowledge, on the stuff that you're able to write down. So this is a little bit more fluid, it seems like, in a sense that it can, in real time, update its knowledge base and its weights and its abilities. So yeah, can we talk about this paper? Because it seems kind of like a, like a big deal, right? Yeah, this paper is even scarier. I haven't done my homework on that one. Oh, no. Joe, if you're Wes, you have, go for it. I'm happy to talk about this paper. Yeah. So this paper is using a very similar idea as the previous one we were discussing. It's like you're asking the model to give you some feedback or some training suggestions. And then as you said, Wes, then you're actually going and doing fine tuning on the model weights themselves. And then if you're thinking about RL training as needing a verifier or a reward signal, your reward is after training the model, like modifying its weights, you run some tests on it that are specific to the thing you trained it on, and you see how well it does. And then that is the reward that comes back to the RL side. So you're sort of training two things at once. You're training the weights of the model on some example, some new data that you want it to understand, and you're training the original model on how well it can suggest training parameters and examples, which is pretty crazy. It's like you're training a model to understand how to train another model. Yeah. And also, you and I talked about this before, Wes, but that's a long cycle. I mean, even though they're using a very small model, I think it's a billion parameters, which is relatively small. It still takes them something like 60 seconds to do the LoRa-based training update and then run through a couple tests, which is a long time to wait for an RL reward signal. Yeah. Yeah. I missed that part of it, but it sounds like the amount of compute is fairly large to do that even on small models. Yeah. Yeah. Yeah. And you can imagine if it was a large model, you might be waiting minutes to get an update, even for training one new data item. So this is a long cycle. The other thing that's scary about this method is it sort of suggests why only do a LoRa update? And notice they throw away the update after they're done, right? They really just want the reward signal. But after you're done doing many such samples, you could collect all of those and use it as a training set if you want to update another model. And furthermore, you could start asking the first model not just for hyperparameters for doing a small training run, but maybe changes to the model itself. Like, can you suggest how many layers should I have or how wide should the layers be? Or maybe we're going to modify the attention mechanism. God knows what changes you want to make to the transformer stack itself, which is something normally that an RL researcher would do or an ML researcher. But you could ask the original model those kind of questions if you wanted to. And now you'd have an even longer period before you got the feedback. Yeah. Yeah. And one of the researchers I noticed on Twitter, and I think you had it in your notes too, did say that the final idea, yeah, is like the teacher and the student model. So you separate them out. All that data is used to train a model that's better and better at suggesting stuff. And then, yeah, I mean, I'm seeing more and more stuff coming out that's similar to Alpha Evolve, right? So the idea is you have these large language models as kind of like the pilots, and then you have like a various scaffolding around it. The model throws out a bunch of, I mean, the trick seems to be how well we can gauge, how well we can evaluate the outputs if we're able to test. Or the Darwin-Godel machine is kind of the same thing, right? So it's able to improve its own abilities to code. So it's like if you're able to evaluate the outputs and then do some sort of that evolutionary tree search, like, oh, this cluster or this lineage of ideas seems to be working really well, let's think through that lineage and continue, like, that seems to be working incredibly well if we're able to evaluate the final output. I just saw another paper where they did the same thing with, it's playing Settlers of Catan. Oh, no. Yeah. And it got pretty good. And it's testing a whole bunch of different stuff. It's doing research online, finding strategies. And it's like, you can probably apply this to, I mean, a lot of things, certainly. And we're probably going to see more and more, just not copy and paste exactly, but they're going to take that idea and just apply it to so many things. I think Dr. Jim Phan and the NVIDIA team were one of the first people that I saw doing it with NVIDIA's Voyager and their Eureka and stuff like that. I mean, when I first saw it, I was like, wait, this thing is improving its ability to train these robots in a simulation just by sampling a bunch of answers from GPT-4 was at the time. And I remember at the end of the paper, they're saying how, like, as the task difficulty gets better, not only does it get better than humans on some of them, but also there's like this divergence between the ideas that humans come up with and what the model comes up with. So it's almost like these novel approaches that we can't necessarily think through, you know, come up with. I was like, man, you know, if nothing stops this progress, did this seem like there's going to be a lot of very interesting applications. Anyways, so the, the anthropic papers talk about how close they are to being able to automate the work of an ML researcher. And I know OpenAI has also mentioned that same metric because what you're hinting at is if they can demonstrate a system that can handle those kinds of tasks, the tasks done by an average ML researcher on an average day, then you would get a sort of takeoff where their automated systems would augment the effort of their own team members. And you would, unless, like you said, unless you top out somewhere, unless there's some diminishing return, you would just get this ramp and the systems would just keep improving. You know, everyone's left the building. Yeah, it's absolutely. Yeah, the, the OpenAI, they have their ML paper bench. I think it is that literally like, can they replicate? So if you give them a, some PhD paper about a machine learning experiment, can they replicate the code base and run that experiment, confirm it? And it's like, it's not quite there yet, but man, it seems, you know, getting better and better. So at some point it's going to cross that line. I feel like so, um, this, uh, self-adapting language models paper is a definite step in that direction. Mm-hmm. Yeah. Yeah. This is interesting. And, and their suggestions at the end are sort of hinting, like if they do another year or two of work, there'll be another big step in that direction. Yeah. How, how far away do you think we are from actually, uh, being able to handle average tasks for an ML researcher and their team? So that's, that's a very, I mean, I have no idea. Obviously I, I deferred to you guys about these things, but I mean, the point is, I think that for everybody listening, everybody, um, on the live chat right now. So, because you guys call out BS when you hear it. So if somebody has some crazy idea about where it's going to go, what it's going to take, you know, you guys are like, nope, here's why. And you guys, you have very good explanations for, for, for it. But I mean, as, as, as we're hearing now, um, the idea of a, some sort of a takeoff once AI research is automated is not some crazy scientific pipe dream. It's, it, you know what I mean? These papers are beginning to suggest that we might be getting closer to it. I think I read, uh, uh, Sam Altman's, um, recent paper, the gradual singularity or gentle singularity. Yeah. Gentle. Yeah. The nice. The nice one. Yeah. The friendly singularity. Yeah, exactly. Here. Back better than ever rebooted. Start the branding early. You got to like, yeah, it's fine. It's a, yeah, no, but one phrase that kind of, uh, uh, jumped out at me, he, he called it. We're at the larval stages of, you know, recursively self-improving. That was so unfortunate when we used larval. I'm like, oh God. I mean, I used to play Starcraft, so I'm thinking of Zerg in my mind. Oh yeah, that's right. Are you saying we're going to get infested now? What's going on? But I think it's kind of a neat way of looking at it. Cause it's like, no, we're not there yet, but we're seeing a lot of these things that seem like if we keep pulling at that thread, it's going to unravel. We're going to, you know, with alpha evolve, all this other stuff, it, it, it's all kind of, we're approaching that recursive self-improving AI and yeah, maybe we'll hit some limit, right? We don't know. Maybe. Well, even, even assuming that you do, uh, see diminishing returns, even assuming that it's not an exponential, it's an S curve. That's, that's okay. An S curve can still take you a long way. And when you do hit the limit at the top, it can reveal other potential improvements. Like we thought we were going to top out with just training of these large models, right? Everyone's freaked out that GPT five isn't, you know, imminent and so much better. And the same thing with cloud four and so on, not to mention meta, uh, four BM off, but, you know, we had test time compute inference time, right? Came in and it was another scaling sort of environment and we're seeing huge improvements there still. And I assume that will top out as well. That at some point, if you give the model, uh, you know, an hour or two to work on something, you won't get any big improvement over giving it a half an hour, say, I don't know what the exact limit is. That's okay. It's another, uh, independent way of scaling the capabilities of the model. And this sort of thing that we're discussing now might be a third or I don't know what, what number we're on now, but another way to get interesting improvements. And even if it tops out, it'll probably reveal other ways. Right. And there's, you know, law of diminishing returns and bottlenecks will always appear in some different way and we'll figure out ways to improve upon it. I think what a lot of people are doing too is trying to sell this, um, panacea of, oh, we're going to hit this point where there's going to be no limitations ever for humanity. And we just put on our 3d glasses, eat popcorn and just let the AGI take off. Um, and for this whole, um, ML researcher, automate ML researcher, I use these models for me, an indicator would be, okay, are these AI labs now slowing down hiring or are they maybe paying less for certain roles? Cause I mean, you're paying a million, 2 million, $3 million a year for some of these researchers. And that could be money that could be going somewhere else in your company. Um, and you know, it might not necessarily work out that, you know, there could be the, uh, Dr. Mike who super appreciate. He could say also, well, no, Jordan, what could happen is now you want to hire more AI engineers and use these models to make them be even more efficient or something. And I think the contrarian would say, I don't know if we want a 50 X or code base to make things more complicated here. We'll be throwing more bodies at the situation is not the answer. So interesting to see all that, all that will shake out. Um, I think we're hitting our one 15, one 20 limit and we're going to move on to the part two. Let's do it. Yes. So, um, we're going to continue this, uh, conversation on the SVIC podcast. So, uh, it's, it's still going to be the three of us. We're just continuing it. We're going to do a little bit more Q and a, I just wanted to limit it, uh, on the side, the live stream to a minute, 20 or thereabouts. So everybody on a live chat, you don't have to do anything. We should, you know, Google willing, uh, YouTube willing, everything should like, yeah. Yeah. Read our, yeah. Oh, oh yeah. No, I said that there's your chance of nothing going wrong so far. It's been surprisingly smooth. I mean, we had some sound issues in the beginning, but sounds like we're pretty good. So everybody hang on for just a second. We're going to redirect this thing and we're going to continue because next we have. Maybe, you know what? We, we spent so much time talking about this stuff, which was very interesting, but I think maybe even some of the more interesting stuff is coming up because what we want to talk about is DeepSeq China. It sounds like the new version of DeepSeq might be resembling the Gemini model a lot more because before it resembled open AI. Now it seemed like it's, um, resembling Gemini a lot more. We got a paper where somebody can not confirms, but basically breaks down their approach to train DeepSeq, the GRPO, I believe, versus the, the PPO and, and, um, some interesting findings there. So we definitely want to cover that. Uh, we have that paper. So stay tuned. Don't do, don't go anywhere. We're going to figure out how to do this. Um, so once we redirect, you guys probably start cause you have everything set up. But let me stop recording in Riverside just to make sure I get everything gets uploaded. So stick with me for one second here and then I will join within one or two minutes. As soon as I'm able to exit this recording, I'll join you guys there. Um, we'll meet in the, uh, the link in the calendar hold that I originally sent out. Perfect. That works. Okay. Everybody stay tuned. Nobody leave. We're going to continue this conversation in just one or two minutes and, uh, I'll see you there. All right. All right. I'll see you guys in there.