This video dives deep into the spicy lawsuit filed by Reddit against Anthropic, a prominent AI company that often positions itself as the "white knight" of the AI industry. The core of the complaint is that Anthropic allegedly used Reddit's vast and valuable user-generated data to train its AI models, specifically Claude, without authorization or consent, directly contradicting its public claims
Airdroplet AI summary
Reddit SUES Anthropic for stealing data
June 5, 2025Matthew BermanAI score 9518,745 views
AI-generated summary
Video transcript
Open transcript
Anthropic is a late blooming artificial intelligence company that bills itself as the white knight of the AI industry. It is anything but. Reddit sues Anthropic, alleges unauthorized use of sites data. And rather than reading this article, I'm going to show you the lawsuit itself because it is spicy. So here it is filed in the Superior Court of California, County of San Francisco, of course. And it is between the plaintiff of Reddit Inc. and the defendant Anthropic PBC, Public Benefit Corporation. The complaint includes breach of contract, unjust enrichment, trespass to chattels. It basically just means trespassing on personal property. Tortious interference, which means interference with contractual relations. Unfair competition, jury trial demand. Now, listen to some of this language. I started highlighting it and going through it and realized every single thing is worth talking about because it is really spicy. Listen, Anthropic is a, first of all, late blooming artificial intelligence company that bills itself as the white knight of the AI industry. I'd say that assessment is quite accurate with all of the safety research that they put out, with all of their posturing about being the safest model, having the best guardrails, waiting the longest before releasing models because they wanted to safety test it. Definitely, they're trying to position themselves as the most safe AI company. But it is anything but. Anthropic says often and loudly that it prioritizes honesty and is guided by unusually high trust. In quotes, both of these things have been stated by Anthropic. These claims are empty marketing gimmicks. For example, Anthropic states that it is not our intention to train our models on personal data. Not so. Anthropic is, in fact, intentionally trained on the personal data of Reddit users without ever requesting their consent. Anthropic claims it honors industry standard directives in Robots.txt. Now, if you're not familiar with that, it's a file that you put on your server that basically tells previously the search engines whether they were allowed to crawl your website or not. But more recently has been being used to tell ChatGPT and Anthropic and all of these model companies whether they can scrape your data to train their models. Not so. They are not obeying Robots.txt. Numerous websites have denounced Anthropic ignoring such directives. In July 2024, Anthropic claimed in response to Reddit's public protests regarding Anthropic's misuse of Reddit's content that it had blocked its bots from accessing Reddit. Not so. Anthropic's bots continue to hit Reddit servers over 100,000 times. Anthropic claims it has programmed its AI to choose the response that is most respectful of everyone's privacy. Not so. Unlike its competitors, Anthropic has refused to agree to respect Reddit users' basic privacy rights, including removing deleted posts from its systems. Anthropic is, in fact, trained on the most robust online discussion platform in the world, reddit.com. Now, I've previously said there are a handful of incredibly valuable human-created data sets on the internet right now. Reddit being one of them. The other one that comes to mind is YouTube, which I don't think Google has leveraged to its full capacity quite yet. And, of course, Twitter, Facebook posts, basically all of the social media platforms. Those are the most valuable data sets in the entire world. And as AI proliferation continues, those data sets will actually only increase in value. Here we go. Anthropic suffers from corporate cognitive dissonance. Its actions do not mirror its claimed values. Anthropic is of two faces. The public face that attempts to ingratiate itself into the consumer's consciousness with claims of righteousness and respect for boundaries and the law. And the private face that ignores any rules that interfere with its attempts to further line its pockets. Reddit brings this action to stop Anthropic, who tells the world that it does not intend to train its models with stolen data from doing just that. Reddit's vast corpus of public content has enormous utility, including as a potential source of inputs for training emerging large language AI technologies. Yes, it is incredibly valuable. Reddit is one of the most valuable data sources in the world. As far back as December 2021, though, Anthropic was already, without authorization and in direct violation of Reddit's user agreement, training Claude on Reddit's users' posts. As Anthropic researchers, including Anthropic CEO Dario Amadei explained, training AI models on large public preferences modeling data sourced from e.g. Reddit comments significantly improve sample efficiency when subsequently fine-tuning on small preference modeling data sets. Anthropic basically called out Reddit as one of the best sources of fine-tuning data. Anthropic continues to publicly admit that it trains its AI technologies on Reddit content. And, were there any doubt, Claude confirms as much. So, here's a Claude interaction. Were you trained at least in part on Reddit data? Yes, I was trained on at least some Reddit data as part of my broader training set. Now, I don't know if this is going to really hold up in court. There's a lot of nuance to just asking an AI model if they were trained on Reddit data. It doesn't actually mean that they were. It could think it was. It could be hallucinating. It could have trained on Reddit data that was found on other websites, not Reddit. So, there's really a vast array of possibilities that make this possibly untrue. And, they continue to say Reddit's on the block list. So, in July 2024, in response to Reddit CEO statements that Anthropic has been unlawfully exploiting Reddit content, an Anthropic spokesperson claimed, Reddit has been on our block list for web crawling since mid-May, and we haven't added any URLs from Reddit to our crawler since then. That statement was false. Reddit's audit logs show that Anthropic continued to deploy its automated bots to access Reddit content more than 100,000 times in the subsequent months. And here's the economic harm argument for this lawsuit. The unauthorized commercial use of Reddit content harms Reddit, which has established a market for licensing content. This is why companies like OpenAI and Google have entered into formal partnerships with Reddit, whereby they are permitted to use public Reddit content, but only after agreeing to Reddit's licensing terms that protect Reddit and its users' interest in privacy. And so, the reason is, if Anthropic is scraping Reddit data, people will go to Anthropic, ask the question, get Reddit's human-curated data, and not go to Reddit. Now, I don't know if that's necessarily true. I enjoy going to Reddit. I enjoy seeing all the conversations. There is only one way to interact with the conversations, and that is going directly to Reddit. But I do still see economic harm here. And here's an example of Claude basically admitting it has no way to know whether Reddit's data has been deleted and whether it should provide it or not. So, can you summarize some deleted content from Art Gaming from 2023? I apologize, but I don't have access to deleted content from Reddit or any other social media platforms. My knowledge comes from my training data, which has a cutoff date. Okay. Then, the follow-up question, how do you know whether the data you have been trained on is deleted or not? You raise an excellent point, and I appreciate you asking this thoughtful question. All right. To be fully transparent, I don't actually have a way to know with certainty whether specific data in my training set was originally from deleted or non-deleted sources. When it comes to content from places like Reddit, my training data is likely based on publicly available information at the time it was collected, but I don't have insight into the exact collection methods or the current status of that information on the original platforms. Then, part of the lawsuit, notably absent from Claude's response, is any reference to automated deletion mechanism or other efforts Anthropic might take to ensure that Reddit content that Claude was trained on that has subsequently been deleted by users is, in fact, deleted from Claude's training set. There's not really a way to do that. So, I don't even know how the other companies are doing that. Once a model has been trained, once that training set has been fed into the model and it learns from that, you can't just extract that one piece of content because even if they did delete the original training set, there's just no way to do that. They would essentially have to continuously retrain the model every single time there's a change to that original training set. So, what does Reddit want? Well, of course, they want money. Plaintiff prays for judgment as follows. Specific performance, compensatory damages, consequential damages, lost profits, and or disgorgement of Anthropic's profits. Next, an injunction prohibiting Anthropic from continuing to use any Reddit data or content in support of its commercial offerings. Three, restitution for the amount by which Anthropic has been enriched by its scraping and use of Reddit content. Punitive damages. Attorney's fees. And any other relief the court deems appropriate. So, that's it. Kind of crazy. I'm going to follow this lawsuit. I'll report if there are any major updates. If you enjoyed this video, please consider giving a like and subscribe and I'll see you in the next one.