this post was submitted on 31 Jul 2026
418 points (99.5% liked)
Technology
86735 readers
2912 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Why would AI want a continuous deal with Reddit? Don’t they get all the data they need the first time? I doubt the new content is worth as much as the previous deal… maybe I don’t understand what these deals are for.
Half the joke is that Reddit was ground zero for AI slop even before AI had gone mainstream.
The company got harvested back before the AI firms were overly worried with cross-contamination.
Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they're still necessary. Which will leave everybody in a bind that is very funny.
Why do they need a deal? Can't they just steal it like the rest of their training data?
The API restrictions and login requirements are meant to make scraping hard enough to make a deal worthwhile.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Other than that, probably it's a licensing agreement that makes AI trainers keep paying if they still using that dataset.