this post was submitted on 23 Sep 2026
528 points (97.3% liked)

Technology

88252 readers
6365 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] Clearwater@lemmy.world 15 points 3 days ago (1 children)

I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser's user agent and those just do whatever they want.

[–] echolalia@lemmy.ml 11 points 3 days ago* (last edited 3 days ago) (2 children)

Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.

I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can't tell who is scraping. They don't scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can't tell who they are, though. How do they get so many residential IPs?

I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.

[–] jsproc@lemmy.world 6 points 2 days ago

It is a business model. People install free apps on their phone. These apps generate their income by acting as residential proxy for scrapers. They form a botnet of millions of ips, actual phones, without the owners knowing it.

[–] Clearwater@lemmy.world 3 points 3 days ago

I'll have to double check my stats later. Haven't looked at it for a while.

I certainly don't recall them ignoring robots.txt, but I also don't remember seeing Meta in my dash at all, so it's entirely possible they, for wherever reason, never found my site.