OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole From Us

ForgottenFlux@lemmy.world · 1 month ago

OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole From Us

Nightwatch Admin@feddit.nl · 1 month ago

It is effing hilarious. First, OpenAI & friends steal creative works to “train” their LLMs. Then they are insanely hyped for what amounts to glorified statistics, get “valued” at insane amounts while burning money faster than a Californian forest fire. Then, a competitor appears that has the same evil energy but slightly better statistics… bam. A trillion of “value” just evaporates as if it never existed.
And then suddenly people are complaining that DeepSuck is “not privacy friendly” and stealing from OpenAI. Hahaha. Fuck this timeline.

Sanctus@lemmy.world · 1 month ago

It never did exist. This is the problem with the stock market.

Ulrich@feddit.org · 1 month ago

That’s why “value” is in quotes. It’s not that it didn’t exist, is just that it’s purely speculative.

Hell Nvidia’s stock plummeted as well, which makes no sense at all, considering Deepseek needs the same hardware as ChatGPT.

Stock investing is just gambling on whatever is public opinion, which is notoriously difficult because people are largely dumb and irrational.

Pasta Dental@sh.itjust.works · 1 month ago

Hell Nvidia’s stock plummeted as well, which makes no sense at all, considering Deepseek needs the same hardware as ChatGPT.

It’s the same hardware, the problem for them is that deepseek found a way to train their AI for much cheaper using a lot less than the hundreds of thousands of GPUs from Nvidia that openai, meta, xAi, anthropic etc. uses

Ulrich@feddit.org · edit-2 1 month ago

The way they found to train their AI cheaper isn’t novel, they just stole it from OpenAI (not that I care). They still need GPUs to process the prompts and generate the responses.

cygnus@lemmy.ca · 1 month ago

Hell Nvidia’s stock plummeted as well, which makes no sense at all, considering Deepseek needs the same hardware as ChatGPT.

Common wisdom said that these models need CUDA to run properly, and DeepSeek doesn’t.

tabular@lemmy.world · 1 month ago

CUDA being taken down a peg is the best part for me. Fuck proprietary APIs.

Fushuan [he/him]@lemm.ee · 1 month ago

They replaced it with a lower level nvidia exclusive proprietary API though.

People are really misunderstanding what has happened.

tabular@lemmy.world · 1 month ago

That’s a damn shame.

Ulrich@feddit.org · 1 month ago

Sure but Nvidia still makes the GPUs needed to run them. And AMD is not really competitive in the commercial GPU market.

ryper@lemmy.ca · 1 month ago

AMD apparently has the 7900 XTX outperforming the 4090 in Deepseek.

Ulrich@feddit.org · 1 month ago

Those aren’t commercial GPUs though. These are:

https://developer.nvidia.com/blog/introducing-hgx-a100-most-powerful-accelerated-server-platform-for-ai-hpc/

Sanctus@lemmy.world · 1 month ago

Someone should just an make AiPU. I’m tired of all GPUs being priced exorbitantly.

3DMVR@lemm.ee · edit-2 1 month ago

they need less powerful and less hardware in general tho, they acted like they needed more

humanspiral@lemmy.ca · 1 month ago

Chinese GPUs are not far behind in gflops. Nvidia advantage is CUDA, drivers, interconnection clusters.

AFAIU, deepseek did use cuda.

In general, computing advances have rarely resulted in using half the computers, though I could be wrong at the datacenter/hosting level at the maturity stage.

Fushuan [he/him]@lemm.ee · 1 month ago

Not cuda, but a lower level nvidia proprietary API, your point still stands though.

Alph4d0g@discuss.tchncs.de · 1 month ago

“valuation” I suppose. The “value” that we project onto something whether that something has truly earned it.

teft@lemmy.world · 1 month ago

I hear tulip bulbs are a good investment…

criss_cross@lemmy.world · edit-2 1 month ago

Nah bitcoin is the future

Edit: /s I was trying to say bitcoin = tulips

boredtortoise@lemm.ee · 1 month ago

Capitalism basics, competition of exploitation

Xanthobilly@lemmy.world · 1 month ago

You know what else isn’t privacy friendly? Like all of social media.

Asafum@feddit.nl · edit-2 1 month ago

You can also just run deepseek locally if you are really concerned about privacy. I did it on my 4070ti with the 14b distillation last night. There’s a reddit thread floating around that described how to do with with ollama and a chatbot program.

Nightwatch Admin@feddit.nl · 1 month ago

That is true, and running locally is better in that respect. My point was more that privacy was hardly ever an issue until suddenly now.

Asafum@feddit.nl · 1 month ago

Absolutely! I was just expanding on what you said for others who come across the thread :)

JoeKrogan@lemmy.world · 1 month ago

Wasn’t zuck the cuck saying “privacy is dead” a few years ago 🙄

NielsBohron@lemmy.world · edit-2 1 month ago

I’m an AI/comp-sci novice, so forgive me if this is a dumb question, but does running the program locally allow you to better control the information that it trains on? I’m a college chemistry instructor that has to write lots of curriculum, assingments and lab protocols; if I ran deepseeks locally and fed it all my chemistry textbooks and previous syllabi and assignments, would I get better results when asking it to write a lab procedure? And could I then train it to cite specific sources when it does so?

WhyJiffie@sh.itjust.works · 1 month ago

but does running the program locally allow you to better control the information that it trains on?

in a sense: if you don’t let it connect to the internet, it won’t be able to take your data to the creators

Asafum@feddit.nl · 1 month ago

I’m not all that knowledgeable either lol it is my understanding though that what you download, the “model,” is the results of their training. You would need some other way to train it. I’m not sure how you would go about doing that though. The model is essentially the “product” that is created from the training.

Sgt_choke_n_stroke@lemmy.world · 1 month ago

chingadera@lemmy.world · 1 month ago

chingadera@lemmy.world · 1 month ago

It just gets better and better y’all.

https://www.theregister.com/2025/01/30/deepseek_database_left_open/

owenfromcanada@lemmy.world · 1 month ago

MysticKetchup@lemmy.world · 1 month ago

I feel like I didn’t appreciate this movie enough when I first watched it but it only gets better as I get older

ouRKaoS@lemmy.today · 1 month ago

“Now” is always a good time to rewatch it & get more out of it!

just_another_person@lemmy.world · 1 month ago

It’s a true comedy that still holds up. I honestly thought for years that Mel Brooks had something to do with it, but he didn’t. It’s so well crafted that there are many layers to it that you can’t even grasp when watching as a child. Seeing it as an adult just open your eyes to how amazingly well done it was.

I could do without the whole Billy Crystalizing of large portions of it though.

owenfromcanada@lemmy.world · 1 month ago

I always thought Rob Reiner had a similar sense of humor to Mel Brooks. And I liked Billy Crystal in it, it kept that section of the movie from feeling too heavy, though I get it’s not everyone’s thing.

For anyone who hasn’t read it, the book is fantastic as well, and helped me appreciate the movie even more (it’s probably one of the best film adaptations of a book ever, IMO). The humor and wit of William Goldman was captured expertly in the movie.

atrielienz@lemmy.world · 1 month ago

I didn’t realize it was a book. Guess I’ll be searching that out.

just_another_person@lemmy.world · 1 month ago

👏👏👏👏👏

A_A@lemmy.world · 1 month ago

x00z@lemmy.world · 1 month ago

Tamaleeeeeeeeesssssss

hot hot hot hot tamaleeeeeeeees

maplebar@lemmy.world · 1 month ago

If these guys thought they could out-bootleg the fucking Chinese then I have an unlicensed t-shirt of Nicky Mouse with their name on it.

sunzu2@thebrainbin.org · 1 month ago

The thing is chinese did not just bootleg… they took what was out there and made it better.

Their shit is now likely objectively “better” (TBD tho we need sometime)… American parasites in shambles asking Daddy sam to intervene after they already block nvidia GPUs and shit.

Still got cucked and now crying about it to the world. Pathetic.

atrielienz@lemmy.world · 1 month ago

They also already rolled back Biden admin’s order for AI protections. So they don’t even have the benefit of those. There’s supposedly a Trump admin AI order now in place but it doesn’t have the same scope at all. So Altman and pals may just be SOL. There’s no regulatory body to tell except the courts and China literally doesn’t care about those.

InFerNo@lemmy.ml · 1 month ago

Now I’m imagining “these guys” are named Nicky Mouse

Cort@lemmy.world · 1 month ago

Oh you want Nickey Mouse, sorry all we have is Mickey Moose.

AbouBenAdhem@lemmy.world · edit-2 1 month ago

DeepSeek’s specific trained model is immaterial—they could take it down tomorrow and never provide access again, and the damage to OpenAI’s business would already be done.

DeepSeek’s model is just a proof-of-concept—the point is that any organization with a few million dollars and some (hopefully less-problematical) training data can now make their own model competitive with OpenAI’s.

Zetta@mander.xyz · 1 month ago

Deepseek can’t take down the model, it’s already been published and is mostly open source. Open source llms are the way, fuck closedAI

AbouBenAdhem@lemmy.world · edit-2 1 month ago

Right—by “take it down” I just meant take down online access to their own running instance of it.

devfuuu@lemmy.world · 1 month ago

Imagine if a little bit of those so many millions that so many companies are willing to throw away to the shit ai bubble was actually directed to anything useful.

SlopppyEngineer@lemmy.world · 1 month ago

deleted by creator

BertramDitore@lemm.ee · 1 month ago

Corporate media take note. This is how you do reality-based reporting. None of the both-sides bullshit trying to justify or make excuses, just laughing in the face of absurd hypocrisy. This is a well-respected journalist confronting a truth we can all plainly see. See? The truth doesn’t need to be boring or bland or “balanced” by disingenuous attempts to see the other side.

I will explain what this means in a moment, but first: Hahahahahahahahahahahahahahahaha hahahhahahahahahahahahahahaha. It is, as many have already pointed out, incredibly ironic that OpenAI, a company that has been obtaining large amounts of data from all of humankind largely in an “unauthorized manner,” and, in some cases, in violation of the terms of service of those from whom they have been taking from, is now complaining about the very practices by which it has built its company.

0x0@programming.dev · 1 month ago

Good that 404 are unafraid of tackling issues, but tbh i find the “hahaha” unprofessional and dispense with the informal tone in news.

BertramDitore@lemm.ee · 1 month ago

I definitely understand that reaction. It does give off a whiff of unprofessionalism, but their reporting is so consistently solid that I’m willing to give them the space to be a little more human than other journalists. If it ever got in the way of their actual journalism I’d say they should quit it, but that hasn’t happened so far.

Fushuan [he/him]@lemm.ee · 1 month ago

It’s just… So deserved, you know? Sometimes you can’t but laugh in the face of such karma and fucking irony.

Rooty@lemmy.world · 1 month ago

I love how die hard free market defenders turn into fuming protectionists the second their hegemony is threatened.

CitizenKong@lemmy.world · 1 month ago

Tale as old as capitalism.

humble peat digger@lemm.ee · 1 month ago

Thank you China.
No for real - it’s either EU or frigging china that helps us with these oligarch overlords

Petter1@lemm.ee · 1 month ago

EU is in best way to become a group of dictators having billionaires in their asses as well…

https://www.politico.eu/article/mapped-europe-far-right-government-power-politics-eu-italy-finalnd-hungary-parties-elections-polling/

rimjob_rainer@discuss.tchncs.de · edit-2 1 month ago

deleted by creator

mArc@lemmy.sdf.org · 1 month ago

the Chinese realised OpenAI forgot to open source their model and methodology so they just open sourced it for them 😂

PrivacyDingus@lemmy.world · 1 month ago

ZILtoid1991@lemmy.world · 1 month ago

Intellectual property theft for me but not for thee!

dogslayeggs@lemmy.world · 1 month ago

Regardless of how OpenAI procured their data, I’m absolutely shocked that a company from China would obtain data unauthorized from a company in another country.

Critical_Thinker@lemm.ee · 1 month ago

It’s a shame that you can’t copyright the output of AI, isn’t it?

ZILtoid1991@lemmy.world · 1 month ago

Trump executive order on the copyrightability of AI output in 3…

vrighter@discuss.tchncs.de · 1 month ago

so? it won’t have any effect on china, because last i checked, us laws apply only in the us

TipRing@lemmy.world · 1 month ago

No honor among thieves.

AwesomeLowlander@sh.itjust.works · 1 month ago

There’s plenty of honor in Deepseek releasing open source.