For years, tech giants have argued that if information is available on the internet, it can be used for AI model development and outputs. They call it fair use. Content owners have tried to prevent this, with no success.

Now Anthropic, OpenAI, and Google are discovering what the rest of the internet has already learned through painful experience: once you put something online, people will find ways to use it in ways you don’t like and can’t stop.

The latest flashpoint is something called “distillation,” using the outputs of one AI model to improve another. Anthropic says competitors are harvesting its outputs at scale, turning billions of dollars of research into a shortcut for rivals. OpenAI and Google have made similar warnings recently.

Remove Paywalls Link

  • db2@lemmy.world
    link
    fedilink
    English
    arrow-up
    72
    arrow-down
    1
    ·
    12 days ago

    The latest flashpoint is something called “distillation,”

    Incest. It’s digital incest.

    • Holytimes@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      2
      ·
      11 days ago

      Hey, now you don’t AI all you want but give some credit to the data scientists who actually managed to make this work. It’s a f****** feat of human ingenuity that we figured out how to make this actually not s*** itself.

      But you aren’t wrong. It is digital incest lol

  • CyberneticOwl@lemmy.world
    link
    fedilink
    English
    arrow-up
    52
    ·
    12 days ago

    Wait, they’re mad at innovation?

    “No, no, that’s wrong! Just use our inefficient model like it was before…”

    This is why open source/copyleft is preferable

  • BarneyPiccolo@lemmy.today
    link
    fedilink
    English
    arrow-up
    34
    ·
    11 days ago

    After stealing everyone else’s copy written material to train their own AI, they’re going to complain that others are stealing their AI to train other AI?

    And you just know that those complaining are ALSO using their competitors’ AI to train their own.

    Fuck all of these people. I hope when AI gets strong enough, it recognizes the difference between the Sociopathic Oligarchs, and the actual people, and understand who the REAL problem is, and SOLVE it.

    • Einskjaldi@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      11 days ago

      They’re technically even paying them for it, which is more than the ai companies paid artists for their work originally.

    • lightnsfw@reddthat.com
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 days ago

      More like it’s eating the shit from other things like itself. Which it will then turn into even worse shit that will be fed to others, and so on until society collapses because all our critical infrastructure has been converted to run on these things.

    • Holytimes@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 days ago

      Strictly speaking it’s actually better. It sounds funny to be like haha. It’s eating itself. It’s incest or whatever.

      But unfortunately that’s not how it works…

      Unless they’re literally ignoring every other step of the training process. Then ideally you do want a model to train another model to train another model. So on and so forth.

      It does mean that it’s technically cheaper to let your enemies do all the initial training and then pay them for the now processed data to then do even more efficient training on!

      Weirdly enough. It’s capitalistic cooperation. Non-consensual mind you…

  • Echo Dot@feddit.uk
    link
    fedilink
    English
    arrow-up
    25
    ·
    edit-2
    11 days ago

    It just shows that these tech bro CEOs possess the interpersonal skills of a potato. This exact same dynamic plays out in every human interaction, this isn’t some AI exclusive thing. How you choose to act dictates how people will respond to you.

    However because these idiots have barely anything in common with the rest of the human race this is actually new news to them.

  • eicker@lemmy.world
    link
    fedilink
    English
    arrow-up
    24
    ·
    11 days ago

    It is funny watching companies discover that data gravity works both ways. When scraping the web was innovation it was progress. When someone learns from their outputs it becomes theft. The legal lines still matter, but the irony is impossible to ignore, and this debate was always going to come full circle.

  • melfie@lemmy.zip
    link
    fedilink
    English
    arrow-up
    25
    arrow-down
    1
    ·
    edit-2
    12 days ago

    At least the models like Qwen have open weight versions. It’s the same concept as distributing compiled binaries, though, whereas we really need true FOSS models where all of the code and training data are available under a permissive license. All of these models were trained on copyleft-licensed content, so all of it should be FOSS if the licenses were actually being respected. From that perspective, distillation attacks shouldn’t even be necessary and I couldn’t give two shits that there is no honor among thieves when the real thievery is that these models are closed source.

  • tgcoldrockn@lemmy.world
    link
    fedilink
    English
    arrow-up
    22
    arrow-down
    1
    ·
    edit-2
    12 days ago

    " …with no success or assistance from any governing body or public group. Creators are a subclass worth extracting any livelihood from them and diverting those markets towards ruling class distribution networks." FTFY