For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”
Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.
This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”
WOW. Holy shit.
I wish that were my reaction to this. Instead, it’s a “DUH. No shit.”
I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM…
Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks
on the one hand, it makes sense if your goal is to train AI to recognize child porn as a simple binary state (bool isCP). Social media sites used to have humans looking at that stuff moderating from afar and it really takes a horrendous toll on their employees.
On the other hand, Elon has repeatedly shown he refuses to censor child porn. They didn’t train it to stop making kiddie porn. They trained it to create more kiddie porn. And that’s why he’s rich. Elon won’t say no. He doesn’t care.
The detail about hash values is really important. The FBI maintains a database of known CSAM. Presumably, hers is in that database, hence the hash values. While not everyone has access to that DB, Xitter/etc does. There is no ambiguity of anything on that list; there’s also no need for any human to review. It’s already been confirmed.
While I’m not sure there’s any case law about it, I would be amazed if using that to train generative AI (except POSSIBLY as content to block) was treated as anything other than possession, distribution, and maybe even production of CSAM.
Proving it might be difficult without full discovery, and AI is infamous for the massive corpus of training data. However, AI doesn’t always generate truly unique works. Go to any image generator and prompt for a video game plumber, and you’ll see an unmistakable image of Mario. It’s possible that they can find a prompt that generates results close enough to her images.
Prompting AI for a video game plumber is most likely going to result in a Mario like entity just because it’s the most common example.
That’s a poor example of what your talking about.
You need someone niche and narrow that has a much smaller sample size. It would be more like asking for a generic description of one persons fursona with out naming it. And the model spitting out a almost 1:1 copy of a real preexisting art work on the fursona. Because the model only has that one picture to base things off of. Which is a real problem.
It’s also the best way to tell if a model was trained on something specific. General prompts aren’t going to get you anything beyond just the fact that yes. X thing is popular enough that everyone and their grandmother creates content on it.
Unfortunately that’s not really how the actual data is stored. Functionally there is no CSAM at all in its “databanks”. Once Info goes through training what comes out the other side is just a mass of goop.
It’s like if you took an entire cow ran it though a meat grinder. Then demanded that you remove only the chuck from the resulting ground beef.
It’s not physically possible.
You can demand it’s retrained entirely from the ground up with vetted data. To produce a higher quality clean dataset. And I would agree that is what should be done.
But you can’t unground the beef
Do you really believe they just throw all the training data away after use?
I don’t believe ANY AI company using internet data to train AI has gone through all of the data to insure it’s not absolute shite.
That seems pretty obvious when you realize that AI is stupid. Most people are stupid, AI is stupid. Tons of pedophiles, AI makes kiddy porn. Millions of racists, AI is racist.
Once I realized this, AI made sense. Checks out.
I’m morbidly curious about the engineers Musk hires, because overwhelmingly the most intelligent engineers and scientists I know wouldn’t dream of applying to work at one of his companies. Even the ones that don’t care about the Nazi shit have still read the many reports about how badly he treats his staff.
i think part of the answer to your curiosity are IT layoffs. job market is shit right now. even though overwhelming majority of people would not work under these conditions if they could choose freely, when they have a mortgage and a family to take care of and they’ve been out of work for long enough the moral concerns and dignity can quickly be overriden by more pressing needs…
Working at spaceX is appealing, because there are not many opportunities to work at the frontier of the industry.
I can’t imagine any 5 year old wanting to work at the boaring company.
What? That’s not even remotely how that works.
A human being is also shaped by the information they’ve encountered throughout their life. You’ve encountered racism, stupidity, misinformation, and countless other forms of bad information. That doesn’t mean you automatically become racist, stupid, or incapable of distinguishing good information from bad information.
AI works similarly in the sense that its training data contains an enormous amount of contradictory, inaccurate, biased, and outright terrible information. The mere presence of that information in the training data does not logically imply that the resulting model possesses those characteristics.
And AI isn’t “dumb” or “smart” in the way you’re describing. Those are human cognitive attributes. The relevant question is what a model can actually do, what patterns it has learned, and how reliably it performs a given task.
You don’t like AI. Fine. But you clearly don’t understand how it works, and this is an unfortunately constant problem on this platform. People make an assumption about AI, then present that assumption as though they’ve discovered some fundamental truth about the technology.
Misinformation is misinformation. It doesn’t become correct because you think the conclusion is morally justified. You didn’t figure out anything about AI here. You made an assumption and called it true.
All, I’m saying if you’re going to critique something, at least know what you’re critiquing or criticizing.
Also before anybody makes any more assumptions, what happen here is absolutely mind-blowingly horrific. Using csam to train anything is just terrible and exploitive. I’m very sorry for doe for having to continuously go through this.
Also also the people responsible for this should be charged with possession of CSAM and prosecuted in a court of law.
Disclaimer: Disagreeing with someone’s argument, pointing out flaws in their reasoning, or interpreting their position differently does not automatically make it a straw man. For example, if I say, “I think we should reduce military spending,” and you respond, “You think we should completely eliminate the military,” you have created a straw man. You are arguing against a position I never actually took.
The mere presence of that information in the training data does not logically imply that the resulting model possesses those characteristics.
Recurring studies do tend to show that AI does indeed posses such characteristics though.
This source doesn’t establish what you think it establishes.
First, it’s a September 2021 infographic, updated in February 2024, not some contemporary study demonstrating that “AI in general” is racially prejudiced. More importantly, the subject here is health-care algorithms, and the article is specifically discussing algorithms that were deliberately designed to use race as a variable or that learned disparities from historical health-care data.
In fact, the source explicitly says that these systems can unintentionally increase existing racial biases through the explicit use of race in predicting outcomes and risk. That’s not evidence that AI possesses racial prejudice. It’s evidence that humans designed algorithms using race as a predictive variable, despite race being a poor proxy for genetic differences. The article even gives examples of medical algorithms where researchers subsequently removed race from the calculation.
That’s an extremely important distinction you’re completely glossing over.
If I build an algorithm that says “Black = higher risk” and the algorithm consequently produces a racial disparity, I’ve demonstrated that my algorithm contains a problematic racial assumption. I have not demonstrated that “AI is inherently racist.” Likewise, if an algorithm uses health-care spending as a proxy for how sick someone is, and that proxy reflects existing racial disparities in access to health care, the resulting bias comes from the data and the proxy, not some intrinsic racial prejudice possessed by the AI.
And your source actually undermines the broader claim you’re trying to make. It explicitly discusses AI being used to reduce racial disparities and cites research where algorithmic approaches improved outcomes or reduced unexplained disparities.
So yes, algorithmic bias in health care is a real and well-documented problem. Nobody is disputing that. What you’re doing is taking evidence that specific algorithms can encode or reproduce human biases and extrapolating it into “AI itself is racially prejudiced.”
That’s not what your source says, and it isn’t what the evidence demonstrates.
*ensure. Ensure means ‘to make secure’, and insure means ‘to have an insurance policy’.
But you’re right. Unfortunately, the biggest weakness with AI is that it’s a funhouse mirror. Problem is, none of us are having fun.
amazing… I know of smaller models that specifically avoided that sort of thing because even negative training could go wrong so easily… they could have avoided that thing entirely, but nope…
I could understand if they were training a safety classifier (a model that learns what danger stuff is so it can recognize it when it sees it) but I doubt this is what they were doing here. That and handling such a radioactive dataset makes handing the demon core seem safe.
“You don’t understand! I need my 400 terabytes of child pornography to keep the kids safe!”
I bet you thats the same excuse the CIA used when they made all those honey pots for lower income pedophiles.
for poor pedophiles
I think I know what you mean, but I strongly urge you to rephrase that in the future.
They probably let it just eat the internet without any humans looking at the training data. And yeah - there probably is some CSAM somewhere on the internet. Maybe, they used an agent swarm to specifically search for stuff not yet in the dataset and forgot to a blocklist. It’s not like the tech bros are genuinely careful in what they do. Recklessness seems to be a common trait.
CCCP notified her that
The USSR?
No, this was the CCCP, not the СССР
If it’s using csam images to produce new ones, is there any way to guarantee that any given nude it makes didn’t source csam? Is the whole thing poisoned at this point?
Anything it makes is derived from the whole of its training data.
Yeah. All the results are tainted, even more than they were from the simple fact it was used by loser to creep on women.
You can’t ever guarantee that a neural network isn’t dreaming of digital sheep - neither for an artificial nor a natural one.
What once has been seen can’t be made unseen again. In that, the clankers are like us.I mean, this is the argument for artist’s copywritten works. There is crossover here. The fights over AI are going to be endless.
To answer your question, yes it’s poisoned.
That’s why I don’t understand why anyone with money to invest in ai doesn’t have a lawyer or lawyers waiving red flags about these issues. It’s copyrighted materials being illegally used, output being machine derived and not able to be copyrighted, the use of csam for training models that output pornography, and on and on. Yet billions of dollars are dumped into what looks like a black market with no concern for how these issues will be settled.
They spend a lot of money on lawyers and congressmen to make sure courts don’t rule that way.
at least in the US there is a legitimate argument that it is fair use, like a mosaic.
Most likely from Musk’s personal collection.
They’re using photo checksums to identify known CSAM? As in, you can change a single pixel to fool it? Is that right?
No. You are thinking about cryptographic hashes, there are locality based hashes that give you a kind of similarity metric
That would be very stupid. I’m sure they have smarter algorithms that can handle a picture being resized of cropped.
Smarter, so long as the end user doesn’t do something crazy like adjust hue and a tiny bit of compaction?
I don’t know what they use but i can come up with multiple smarter ways on the spot.
Its probably a multitude of hashes and finding multiple ones in a picture is a red flag.
one i would do is to compare the difference between pixels. You can change the colors all you want. Dark hair versus teeth will always read as opposites per example.
Actually i don’t need to argue that it works… ai generation is always an entirely new picture with new pixels in a set resolution … thats why ai cant do colorisation right.
The fact that the existing hashes where detected in such recreation means the tech is very impressive. I doubt false positives are common with these either.
If Doe is actually recognizable, then chances are high that actual pictures of her or him were used for training the AI. Wouldn’t surprise me. They scrape the internet for everything they can get their digital hands on, regardless of copyright or criminal law. They are bound to find illegal stuff on that track.
The next thing is that there are probably confidential data in the training sets, either exposed by neglient users, or by hackers that breached sites and blackmailed them.
And of course all the copyrighted material they used without permission.
If AI companies would really get sued on those three illegal sources, they could probably close their doors.
Next is a “funny” issue: scraping the web without checking sources will inevitably lead to AI trained on illegal content.
The only way for them to prevent that is to train their AI to recognize illegal content. And the only way to do that is to feed it that content with label.
AI corps would have to pay people to watch children sexual abuse and label the videos. (Not that I have the slightest doubt they would proceed if that was allowing them to keep going rampage on web scraping)
People have that job already. How do you think that websites where you can upload amateur porn, filter and block CSAM? Someone has to view the flagged material.
Sounds like another MAGAt of the year action.
Not big surprise. We live in a “do evil now, ask for forgiveness” later world. I’m getting tired of evil.
This is beyond fucked up. There should be laws in place such that if a company is found to be using CSAM in any way, the company ceo and anyone else involved in the work that uses it is personally criminally liable. No hiding behind a corporate shield with this stuff.
bruh
freaks
that is a great looking inflatable
Look up AI over fitting. Because of the math, every answer can be provided in terms of how much CP data was used in every answer.








