Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

  • Kissaki@programming.dev
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    51 minutes ago

    Noteworthy: They crawled only the default branch HEAD and inlined all source content.

    • The file contents are included inline. The decoded UTF-8 source text is embedded directly in the dataset, so it is fully self-contained — you can start training the moment the download finishes.
    • It reflects the state of GitHub in August 2025. The corpus is a direct crawl of GitHub repositories at their default-branch HEAD, capturing roughly two additional years of open-source code compared to The Stack v2.
  • Kissaki@programming.dev
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    46 minutes ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    crawled directly from GitHub and built to pre-train code LLMs with full-repository context

    Repositories that opted out are removed from the dataset before each patch release.

    “agency”

    Which AI company will not use v1 which has all of the data but will use later patch releases instead which have less data?

  • Kissaki@programming.dev
    link
    fedilink
    English
    arrow-up
    1
    ·
    41 minutes ago

    From the linked webpage readme:

    2.1 Classify. Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.

    From https://www.bigcode-project.org/docs/about/the-stack/:

    v1.1: The three copyleft licenses (MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming languages was increased from 30 to 358 languages. Also opt-out request submitted by 15.11.2022 were excluded from this ersion of the dataset. The resulting near-deduplicated dataset is 6TB in size.

    So MPL/EPL/LGPL are already not part of the dataset.

    So… why were they in there? Was this added for v1.1?

    “one permissive license” - So if my project includes a lib and I include the license file for that for the license notice…?

  • qaz@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    2 hours ago

    I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.

    Surely there has to be a better way to do this (probably the point).

    Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?

    EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership

    EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?

    • Kissaki@programming.dev
      link
      fedilink
      English
      arrow-up
      1
      ·
      47 minutes ago

      with a Markdown list with all your repositories you want removed.

      The repo readme linked FAQ says

      You can choose to request either (1) all repos, or (2) you can specify select repos that you own to be removed.

      so “all of them” should be acceptable

  • Prior_Industry@lemmy.world
    link
    fedilink
    English
    arrow-up
    13
    ·
    8 hours ago

    When I hear this company’s name I can’t get the idea out of my head that it’s related to the face huggers from Alien.

  • purplemonkeymad@programming.dev
    link
    fedilink
    arrow-up
    9
    arrow-down
    1
    ·
    10 hours ago

    Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 hours ago

      The right move is to private all repositories you are uncertain of that you want them to be preserved forever in an AI training set.

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 hours ago

      They also scraped some of my assembly code of which I’m pretty sure the latest version has a major bug. Let’s hope they don’t use these models to program my new PC’s bios.

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      edit-2
      3 hours ago

      No, it seems to only be a subset of public repo’s.
      I have like 65 repo’s and only 13 were scraped. I don’t get why they specifically scraped those though. They don’t have the most stars, they aren’t the oldest or newest, not the ones with the most forks, nor do I see a pattern based on programming language.

  • SpaceCowboy@lemmy.ca
    link
    fedilink
    arrow-up
    52
    arrow-down
    2
    ·
    20 hours ago

    A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”

    Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.

    • FishFace@piefed.social
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      11 hours ago

      Making arguments from popular perception is never very strong. AI is cool if you actually think about it - the capability is incredible. Anything that can produce working code was going to have this ambivalent result where execs pushed it way too hard.

        • NewNewAugustEast@lemmy.zip
          link
          fedilink
          arrow-up
          1
          ·
          edit-2
          4 hours ago

          Simultaneously they have created the part I wanted from Star Trek, while making the worst possible anti star trek a reality.

          E.g.

          I want to be able to ask a computer about history, art, codeing, well anything. And be able to clarify and question and put together new ideas.

          But not by anyone owning that ability or profiting on the labor of others or causing environmental harm.

          We got the cool computer but haven’t achieved the post-scarcity part.

          I don’t think you can have one without the other.

    • FizzyOrange@programming.dev
      link
      fedilink
      arrow-up
      5
      arrow-down
      1
      ·
      13 hours ago

      I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.

      The level of AI we have now was impossible science fiction a decade ago.

  • abacabadabacaba@infosec.pub
    link
    fedilink
    English
    arrow-up
    52
    arrow-down
    2
    ·
    23 hours ago

    The best way to opt out is by not using GitHub. Also opts you out of Copilot and a bunch of other stuff.

    • fruitcantfly@programming.dev
      link
      fedilink
      arrow-up
      12
      ·
      edit-2
      13 hours ago

      Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        3
        ·
        13 hours ago

        I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.

        I like to have the visibility and CI of GitHub. But this sucks ass.

    • Eager Eagle@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      edit-2
      22 hours ago

      note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset

    • DarkCloud@lemmy.world
      link
      fedilink
      arrow-up
      4
      arrow-down
      1
      ·
      edit-2
      22 hours ago

      When I back something up, I save a copy and add “backup” to the name… Because I’m advanced.

      If I’m feeling really good and healthy, I’ll even put it on a usb stick.

  • G_M0N3Y_2503@lemmy.zip
    link
    fedilink
    arrow-up
    12
    arrow-down
    1
    ·
    19 hours ago

    Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.

      • Alex@lemmy.ml
        link
        fedilink
        arrow-up
        7
        ·
        16 hours ago

        Mine are all GPLv3 or forks of other repos and they are listed. I’m sanguine because the license allows for study and if it’s good enough for humans I don’t see why it’s not for clankers.

        • fruitcantfly@programming.dev
          link
          fedilink
          arrow-up
          6
          ·
          13 hours ago

          For me, and a few other projects I checked, it only has non-GPL repos. But it also does not have everything that isn’t GPL, despite the repos being much older than the cut-off date. But it does have repos without a license, which they are simply not allowed to copy.

          I wonder if those repos have been deduplicated, and one of the forks (on some other person’s account) is included instead. Unfortunately you can only search the first 5M records via the website, and I don’t have time to play around with the API at the moment, so I could neither confirm nor deny that possibility

  • chicken@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    27
    ·
    22 hours ago

    A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site

  • goatbeard@beehaw.org
    link
    fedilink
    arrow-up
    9
    ·
    edit-2
    20 hours ago

    Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners