DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
60% Positive
Analyzed from 4558 words in the discussion.
Trending Topics
#books#book#copyright#old#rare#digital#more#physical#archive#copies

Discussion (77 Comments)Read Original on HackerNews
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
Who is we?
> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely
Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.
Yet
As for consumer-grade solutions, look for the Fujitsu SV600.
I have done it at home for my books since the mid-00s.
But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.
IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.
This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.
The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
Because there’d be much less content created in any media to capture in the first place.
An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.
Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.
they should pay the marginal value that the next buyer would buy.
Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?
Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.
Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.
[1]: https://en.wikipedia.org/wiki/Decoupage , since I imagine a number of people won't know what that is.
they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.
having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.
> paying extortion fees to the copyright-mongers
Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s
Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.
And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.
A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.
For whom? is the relevant question
https://www.ala.org/tools/challengesupport/selectionpolicyto...
Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.
What similarities do you see here?
Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.
A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
What guarantee do we have that the book contents will be served unfiltered and unaltered?
I wish...
This is highly disturbing news; is this standard practice? What did Google Books do before?
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
https://skybrian-links.exe.xyz/post/1026
> It is equivalent to book burning in the past. A form of thought control
> > It is equivalent to book burning in the past. A form of thought control
That's only bullshit if you trust AI companies to serve the book contents without alteration.
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.
If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.
Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.
Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!
There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.
For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
What makes you think they will? What would be the incentives for these companies to do so?
1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.
2. Availability via Google Books or similar.
3. Availability via AI model reference.
4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.
I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.
I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.
What if the perception of those books change over time and are considered masterpieces later on?
Moby Dick was out of print when Melville died 1891 and not a huge success until it got a revival in the 1920s
That’s the whole point because the book shredding is already declared legal.
https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act