ES version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
46% Positive
Analyzed from 4994 words in the discussion.
Trending Topics
#books#book#rare#copyright#don#training#companies#copies#more#going

Discussion (89 Comments)Read Original on HackerNews
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
Real humans sharing their unique knowledge, packaged in a book.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
I would be also interested in participating in your costs.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
Thank you.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
Not even the title of one of those rare books?
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
The rare qualifier is used precisely because no reasonable person thinks this is stealing.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
I think the common consensus is that stackoverflow killed stackoverflow, quite a few years before LLMs became entrenched.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
https://www.authorsalliance.org/2025/09/07/the-anthropic-set...
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
2 days ago https://news.ycombinator.com/item?id=49310725
21 days ago https://news.ycombinator.com/item?id=49068738
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
The verbatim content are the words, not the paper.
Books are lost all the time because the last book ended up in a landfill. But if an AI lab digitizes it, now it's stored in an extremely redundant storage lake in a datacenter and the company has huge incentives to make sure they don't ever lose that data.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.