BetaDoggo_

joined 2 years ago

The AI company Perplexity is complaining their bots can't bypass Cloudflare's firewall in c/technology@lemmy.world

[–] BetaDoggo_@lemmy.world 21 points 23 hours ago* (last edited 23 hours ago) (4 children)

Perplexity (an "AI search engine" company with 500 million in funding) can't bypass cloudflare's anti-bot checks. For each search Perplexity scrapes the top results and summarizes them for the user. Cloudflare intentionally blocks perplexity's scrapers because they ignore robots.txt and mimic real users to get around cloudflare's blocking features. Perplexity argues that their scraping is acceptable because it's user initiated.

Personally I think cloudflare is in the right here. The scraped sites get 0 revenue from Perplexity searches (unless the user decides to go through the sources section and click the links) and Perplexity's scraping is unnecessarily traffic intensive since they don't cache the scraped data.

permalink
fedilink
source
context