Would be a terrible shame if lots of people opted out.
If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.
lol my first ever JavaScript project is in the stack. No wonder their models suck so much ass
my commits are the reason we haven’t achieved AGI
Thank you for your service 🫡 please keep committing trash code
My Arduino led project that didn’t work is going to topple the world order.
They also scraped some of my assembly code of which I’m pretty sure the latest version has a major bug. Let’s hope they don’t use these models to program my new PC’s bios.
Same here. Out of all my public facing repos, they grabbed the group projects from when me & friends were first learning to code 👀
A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”
Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.
I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.
The level of AI we have now was impossible science fiction a decade ago.
Making arguments from popular perception is never very strong. AI is cool if you actually think about it - the capability is incredible. Anything that can produce working code was going to have this ambivalent result where execs pushed it way too hard.
Yup, the idea is good. How this this idea is being implemented… not so much.
Simultaneously they have created the part I wanted from Star Trek, while making the worst possible anti star trek a reality.
E.g.
I want to be able to ask a computer about history, art, codeing, well anything. And be able to clarify and question and put together new ideas.
But not by anyone owning that ability or profiting on the labor of others or causing environmental harm.
We got the cool computer but haven’t achieved the post-scarcity part.
I don’t think you can have one without the other.
Post-scarcity isn’t actually possible. Not even in Star Trek this is true. Picard’s family owns a vineyard in France filled with artifacts and antiques. Not everything is fungible. People will desire these non-fungible things. Not everyone that desires these things will be able to have them, because they aren’t fungible. You can’t have a billion people all owning vineyards in France. Some people won’t get everything they want. There will always be scarcity.
Star Trek was made in the 1960’s at the height of the cold war. They didn’t want the show to be about how capitalism was superior to communism, or vice versa. So they side stepped the issue by saying in the future there’s no scarcity, no money, and they live under some ideal future economic model. The show writers don’t know what that ideal economic model is, so it’s deliberately vague and inconsistent. And that’s fine because the show isn’t about economics.
It was wise of the writers of Star Trek to avoid trying to make predictions of that nature. 300 years ago, Wealth of Nations wasn’t yet written, the field of economics didn’t really exist. They believed shiny rocks had intrinsic value. They had some vibes about things like currency devaluation and inflation, but economics was still mostly about acquiring shiney rocks 300 years in the past. So what will economics be like 300 years in the future? None of us know, and certainly writers of a TV show don’t know.
The writers of a TV show saying there will be no scarcity in the future just means they didn’t want to discuss economics. It’s not any kind of prediction about the future. There will always be scarcity.
Doesn’t hugging face do open source AI?
The best way to opt out is by not using GitHub. Also opts you out of Copilot and a bunch of other stuff.
Already did that some time ago but I am still in the dataset.
Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge
I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.
I like to have the visibility and CI of GitHub. But this sucks ass.
note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset
When I back something up, I save a copy and add “backup” to the name… Because I’m advanced.
If I’m feeling really good and healthy, I’ll even put it on a usb stick.
A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site
When I hear this company’s name I can’t get the idea out of my head that it’s related to the face huggers from Alien.
Wonder how many of those repos contain AI-generated code?
if I opt out, will the Roko’s basilisk come after me?
Reported for cognitohazard. Delete this immediately.
(I kid, but someone really did report your comment.)
lmao
I think its hilarious that some people seem to actually take this science fiction variation on Pascal’s wager seriously.
No because your code is shit; you would be helping us reach singularity faster.
Much like… -reads notes- Microsoft is via Github… huh…
Gothub.com is a thing now for OpenBSD’s Game of Trees project.
It’s gothub.org. You linked to some dumb AI startup.
It’s a sure sign of a bubble that a URL typo coincidentally brings one to a completely unrelated page for a completely unknown, in-beta AI startup with a dumb idea.
Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.
Strange because all of my MIT repos are listed - all the GPL ones are excluded.
Mine are all GPLv3 or forks of other repos and they are listed. I’m sanguine because the license allows for study and if it’s good enough for humans I don’t see why it’s not for clankers.
For me, and a few other projects I checked, it only has non-GPL repos. But it also does not have everything that isn’t GPL, despite the repos being much older than the cut-off date. But it does have repos without a license, which they are simply not allowed to copy.
I wonder if those repos have been deduplicated, and one of the forks (on some other person’s account) is included instead. Unfortunately you can only search the first 5M records via the website, and I don’t have time to play around with the API at the moment, so I could neither confirm nor deny that possibility
My list includes GPL licensed repo’s
Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?
The right move is to private all repositories you are uncertain of that you want them to be preserved forever in an AI training set.
We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.
crawled directly from GitHub and built to pre-train code LLMs with full-repository context
Repositories that opted out are removed from the dataset before each patch release.
“agency”
Which AI company will not use v1 which has all of the data but will use later patch releases instead which have less data?
Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners
I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.
Surely there has to be a better way to do this (probably the point).
Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?
EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership
EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?
with a Markdown list with all your repositories you want removed.
The repo readme linked FAQ says
You can choose to request either (1) all repos, or (2) you can specify select repos that you own to be removed.
so “all of them” should be acceptable
tbh i’m thinking this alone isn’t that bad from an archiver/datahoarder perspective
As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.
This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?
IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.
Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks
Internet Archive exists as a reference for your edification on a donation basis; AI companies intend to initiate a top down societal restructuring of jobs and thus access to benefits, paywall access to your own collective information, fund themselves through ouroboros leveraged deals and VC money (value that’s been stolen from the general populace over the years)
deleted by creator
They’re not able to scrape private repos, surely?
Microsoft has access, and everything is for sale. So, maybe? Probably? 🤷
No, it seems to only be a subset of public repo’s.
I have like 65 repo’s and only 13 were scraped. I don’t get why they specifically scraped those though. They don’t have the most stars, they aren’t the oldest or newest, not the ones with the most forks, nor do I see a pattern based on programming language.











