this post was submitted on 21 Jul 2026
229 points (99.1% liked)
Technology
86503 readers
4163 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I'm waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.
This is already happening.
They are paying people bottom market rates to generate unique written content and people are just having AI write it for them.
Some kind of AI ouroboros situation.
(Non)Human centipede.
Reverse centaur centipede?
Supposedly the people paid to rate responses are already using AI
"now you only posted 4 times today Jenny, do you just want to do the minimum?"
Big AI mostly switched to synthetic training data anyway. The books they're digitizing are being used to gather knowledge, not writing styles or logic (mostly).
As in, when you ask ChatGPT how long some book is, it can just go check (if it's in the database). It's also useful if you ask about that book or about knowledge contained in that book. It'll even reference books now (if you demand that in your prompt).
It's not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The "language" part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷
The article quotes a post from ISBNdb saying the issue is model collapse from training on synthetic data.
Yes, what's the issue?
That's like saying, "they had some failure modes from the synthetic data, so they should just obviously stop trying forever."
They'll just fix the edge cases and move on. Like any programming task.
Total collapse is a solution.
It's also unavoidable.