The symbol for Amazon’s VGT3, the Las Vegas facility where it scans book for AI training data.

Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.

A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.

“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.

The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.

In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called “model collapse.”

📖 Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at emanuel@404media.co.

These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.

In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order. 404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.

Another article to provide some clarity on the types of books: https://finance.biggo.com/news/37a2899b-5f1a-4571-9bb5-1156c0ac4605

    • Test_Tickles@lemmy.world
      link
      fedilink
      English
      arrow-up
      16
      ·
      13 days ago

      Look up Google and Project Ocean. Google already did this 2 decades ago. They even went to great lengths to build machines that would very slowly and gently turn pages and non-destructively scan books.
      They already have a digital library of 25 million books just sitting there, ready to be instantly and non-destructively copied infinitely.

      • Mirshe@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        arrow-down
        1
        ·
        13 days ago

        No no, that’s too slow and slow is expensive. Time is money and we have to move faster and break more things in order to disrupt the market. /s

      • ThomasWilliams@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        11 days ago

        But a court ruling said they had to destroy the books otherwise they would be liable for copyright infringement

    • Solrac@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      12 days ago

      You don’t hate them enough. The only response is to do onto them, as they do to us, as they do to these books

  • hard_zero1@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    29
    arrow-down
    1
    ·
    13 days ago

    If this was done by trustworthy organizations as an effort to digitize and preserve all the books, I would appreciate it. Maybe an association of libraries should do that and, to get the cost back, sell the digital versions / ebooks to the AI companies for an additional price). Then, at least, we would not loose the contents of those rare books to the AI conpanies and prevent them from obtaining a monopoly on the data. And each book would only get destroyed once.

    But libraries/bookshops are probably not allowed to sell digital versions, and AI companies are not allowed to use borrowed ebooks.

    • Duamerthrax@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      12 days ago

      That exists already. Archive.org has a digitization service, but then archive.org would put a public copy up and the AIbros wouldn’t be the sole owners of that training data.

      https://digitization.archive.org/

      There’s also methods to digitize books without destroying the book and for the purpose of AI training that should be more then sufficient. AIbros are just so arrogant to think that other people haven’t solved the problem already or that their time is too valuable to be slowed by proper methods.

      https://www.diybookscanner.org/

      • ggtdbz@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        1
        ·
        12 days ago

        I remember looking into this, the number of books I have that I’d like to scan are not enough to justify building one of these unless they can be folded up and passed on.

        Might be best to just get a plexiglass slab with some kind of anti glare covering and try to get flat pages

    • frongt@lemmy.zip
      link
      fedilink
      English
      arrow-up
      10
      arrow-down
      1
      ·
      13 days ago

      Yeah. Libraries got in trouble during COVID for relaxing their ebook borrowing rules. AI companies don’t give a shit and are happy to settle lawsuits for a fraction of the money they take in from investors.

    • distal@lemmy.ml
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 days ago

      No, even this is a very bad idea. Ideally, no books are destroyed. A close relative of mine works as an archivist. They have tons of examples of how digitising and then destroying can create extra burdens. The data on hard drives or flash devices is too fragile.

  • queermunist she/her@lemmy.ml
    link
    fedilink
    English
    arrow-up
    14
    arrow-down
    3
    ·
    13 days ago

    I don’t understand the concept of a “rare” book. Every book should be infinitely reproducible, the fact that they aren’t is a crime.

      • queermunist she/her@lemmy.ml
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        5
        ·
        edit-2
        12 days ago

        The text is what makes it valuable though? I guess there’s some value in, like, the actual literal physical copy too, but that has nothing to do with rarity. A bible can have historical value because of who owned it, but that’d be strange to describe as “rare” I think.

          • queermunist she/her@lemmy.ml
            link
            fedilink
            English
            arrow-up
            1
            ·
            edit-2
            11 days ago

            A first edition has no value to me. The value comes from the words on the page, not the paper it’s printed on. I could see value in a book that was hand transcribed or something, but the fact that it was the first book out of a printing press is worthless.

            I should say being signed is important, but that’s not really a rarity issue. That’s now a unique historical artifact because it was interacted with by a specific person. That’s basically the same as a bible having historical value because of who owned it.

    • ThomasWilliams@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 days ago

      They only print a few dozen copies and then Google destroys the only one left.

      It’s not a difficult concept.

      • queermunist she/her@lemmy.ml
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        2
        ·
        11 days ago

        It’s not like we’re printing books from old school printers where they had to assemble the letters by hand on a printing plate for each page, or they need scribes to hand transcribe books with a quill. It’s all computer. There’s no reason for any book to be rare.

  • altkey (he\him)@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    9
    ·
    13 days ago

    The symbol for Amazon’s VGT3, the Las Vegas facility where it scans book for AI training data.

    I thought it’s 404media’s obviously satirical preview picture to dab on aibros, but the truth is even more weird

    • zarkanian@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 days ago

      Are you skeptical that the books being destroyed are actually rare? Or are you skeptical of the existence of rare books?

    • Tollana1234567@lemmy.today
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      3
      ·
      edit-2
      12 days ago

      it should be labeled as “limited amount of books in circulation”, not rare. rare means 1 or only handful of compies.

      • zarkanian@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        3
        ·
        11 days ago

        “Rare” means “difficult to acquire”. Under your definition, the only books that would qualify would be ones in museums.

      • distal@lemmy.ml
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        11 days ago

        I know that in disciplines such as math there are very often have books with less than 50 copies in circulation. The LLM companies buy maybe 2 of a given book which are subsequently destroyed with no backup data. You’re now at 48 in circulation, that is a significant reduction. Rinse and repeat a couple of million times and this starts to have a major influence on knowledge.

        What’s perverse about it is that the data from the books is stored extremely badly within the models. So not only are you destroying sources, you are also necessarily falsifying them.

  • BilSabab@lemmy.world
    link
    fedilink
    English
    arrow-up
    5
    arrow-down
    1
    ·
    12 days ago

    local AI companies are the reason Ukrainian book publishing industry experienced an uptick in sales lately. Bros stock up entire catalog at once and no one complains because its a hefty paycheck.

    • KairuByte@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      2
      ·
      11 days ago

      .> Local to… Ukraine? Why does this feel like a conspiracy theory? Ukraine, in the middle of a war, is focusing their efforts on an LLM, destroying books in the process?

      Are you a plant?

      They may very well be doing just that, but it’s really fucking weird to bring it up in a post about Amazon destroying rare books in the US.

  • BillCheddar@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    11 days ago

    They also want the rare books scanned so they can break any/all book ciphers, including those based on rare finds.

  • hard_zero1@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    2
    ·
    13 days ago

    While writing my other comment, I thought of a way that might be a little bit closer to a solution for the AI use of copyrighted material. In Germany, afaik we are allowed to make and keep a copy of copyrighted material obtained in any legal way (even borrowing, streaming, etc.) for strictly private and personal noncommercial use. As a compensation, when buying a device that can store or process data (e.g. USB-stick, phone, computer), a small part of the money goes to copyright collectives (“Verwertungsgesellschaften”). Copyright holders can be members of those and get compensation or their work via them (those collectives are also the entities that sell you a license to e.g. play music at public events).

    I can imagine a similar system for AI: A part of any subscription, fee, etc. for AI could go to those copyright collectives as a compensation for the material used in training. Compared to obtaining a license once for training (like buying physical books), this has the advantage that the compensation would scale with the actual use of the AI and also the revenue of the AI companies.

    I think those systems are far from perfect, but that’s probably a much better approximation to fairness than what we currently have

    • Funkt4st1c@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      13 days ago

      Would make sense if the AIs made any money. This is kind of the same system we currently have while AI burns millions daily. Dollars, tons of coal, gallons of water reserves, creatives’ livelyhoods… take your pick

      • hard_zero1@discuss.tchncs.de
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        13 days ago

        At least some money is currently paid for AI and eventually the business has to become profitable. The environmental problem ist of course not approached by this. For the creatives this would give at least some compensation. Problems due to future (creative) work taken over by AI is also not covered by this; that’s “just” what happens when a new technology arises that obsoletes certain work. My suggestion is trying to copensate for creative work somewhat proportionally to its use.

        • Funkt4st1c@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          12 days ago

          The business has to become profitable… To continue existing. There is absolutely no guarantee of that.

          There is, however, overwhelming evidence that unless a major breakthrough is discovered within the next year, the well will dry up. They simply cannot afford to live off of investor hype and vibes forever.

  • antonim@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    ·
    13 days ago

    So, what is the book that they tracked? What other books were in the order alongside it?

    • Jiral@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      13 days ago

      You are aware that they’d be betraying their source if they published the information that could idenify the seller?

      • antonim@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        1
        ·
        13 days ago

        Perhaps.

        What I actually want to know is the rarity of the books in question.

          • antonim@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            1
            ·
            12 days ago

            Thanks. I can’t help but conclude that the public outcry is exaggerated – the books being bought are evidently selected at random. The chances of the companies buying some particularly rare (and more expensive) book this way seem slim, since they probably don’t want to waste too much money.

            • zarkanian@sh.itjust.works
              link
              fedilink
              English
              arrow-up
              1
              ·
              edit-2
              11 days ago

              Money is no object to AI companies. That’s one of the things that alerted the booksellers, in fact—masses of books were being purchased without regard to price or subject matter.

            • ThomasWilliams@lemmy.world
              link
              fedilink
              English
              arrow-up
              1
              ·
              11 days ago

              This is nonsensical. Anything valuable is copied anyway. What they are destroying is thousands of small books about local histories and technical matters that are being disappeared entirety.

    • roofuskit@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      edit-2
      11 days ago

      What’s their standing?

      Edit: FYI in the United States we have a legal precedent called the first sale doctrine. Once you buy a book it is yours and you have the right to do what you want with it including selling it. And while they probably never anticipated a bunch of super wealthy narscisists buying a bunch of books so they could hoard the knowledge and then destroy them, we have no laws here that could prevent it.

  • melsaskca@lemmy.ca
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    1
    ·
    13 days ago

    I bet there are a whole lot of politicians books mixed up in there. That’s the new payoff method for the ruling class.

  • mr_tyler_durden@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    2
    ·
    11 days ago

    ISBNs or STFU.

    “Rare” means absolutely nothing without context. This is trash reporting, you sent the books, you know what they are, by not publishing the titles/ISBN it proves you know you don’t have a valid case. I’m betting these were mostly/all used books.

    The pearl clutching about this is absurd. So what if they buy 1 of every book to scan?