this post was submitted on 28 Aug 2026
811 points (98.7% liked)

Technology

87624 readers
5081 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

you are viewing a single comment's thread
view the rest of the comments
[–] uriel238@lemmy.blahaj.zone 2 points 8 hours ago

I thought that Google used a similar dataset, possibly borrowing images from the NCVIP in order to create an analytic rule-set for which to omit images from Google Image Search. (A larger, similar rule-set is used to omit legal pornography when safe-search is on.)

Mind you, this was before the hyperscale AI era, when LLMs were things like SIRI and Google Now. And Google search still focused on websearch hits and not AI summaries.

Fun story: This was an area of study of mine during the early 2010s, since every image search engine would filter porn hits whether or not you had safe-search on or off. If it was turned off, porn would be filtered to the end of the list unless the engine decided you were intentionally looking for porn in which case the porn hits would be shown at the top of the list. There was no way to get results that ignored the ID-as-porn status of the images. I would enter ambiguously risqué terms to see how explicit I needed to be before the engine decided I was looking for porn.