Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
43% Positive
Analyzed from 4439 words in the discussion.
Trending Topics
#books#book#companies#copyright#copies#copy#more#rare#destroy#archive
Discussion Sentiment
Analyzed from 4439 words in the discussion.
Trending Topics
Discussion (145 Comments)Read Original on HackerNews
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
They just don't want to pay what the copyright holders want to charge
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
I support Anna's Archive, by the way. Information wants to be free.
https://annas-archive.gl/donate
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
No. They also use lots of other methods to get training data.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
What if people like food more than AI? Have you considered that?
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
Which exists to enrich Disney and other large corps, while they hide behind artists as a human shield.
>Artists need some form of reward.
Right, as do artists who use the work of other artists as their starting point. Copyright holders aren't bill and bob artist, they are massive corporate trolls throwing around the weight of almost 100 years of our cultural heritage, sucking the marrow from its bones.
>On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
The books are getting digitised into a permanent record of all our cultural heritage. It just sucks we don't have control over it. If only there was a way we could get them digitised AND control our cultural heritage. HMMMMMMM.
>Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write.
The incentive to write is being killed by slop groups like 20Booksto50K and Kindle which predate AI by at least a decade. AI just lets them work faster.
>We might end up with all of humanity’s books digitized and accessible for free
Excellent
>But there would be no human writers left.
Unlikely, but there would definitely be no Disneys or Conde Nasts left, which is a massively pro social outcome.
>In a world like that, what motivation would we still have to read?
In a world with all books digitised and accessible to read? A huge huge huge incentive. I already partake if books are too expensive where I am. It would take me the rest of my life to read all the books I already want to read. What kind of inane dribble is the idea that copyright makes it interesting to read? I havent even read all of Howard and he's in the public domain (in cool countries at least)
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
No data => No models => No competition.
> ChatGPT […] originally released on November 30, 2022
https://en.wikipedia.org/wiki/ChatGPT
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
Edit; despite the above, looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form we purchased, therefore it is legal.”
I still dont think there is a requirement to destroy the book, but since there isn’t a reason to store the book and they can’t sell it, they probably just took the cheapest route. It might be worth an argument that they only purchased the right to use the digital copies while the physical copies exist, but I’m in over my head from a copyright standpint
I imagine since the law recently cost one of them truckloads of money for their violations of it?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
-- Thos. Jefferson
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.