Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
59% Positive
Analyzed from 6394 words in the discussion.
Trending Topics
#books#book#copyright#rare#old#physical#companies#should#more#digital
Discussion Sentiment
Analyzed from 6394 words in the discussion.
Trending Topics
Discussion (174 Comments)Read Original on HackerNews
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I'd like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.
You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.
At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.
Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards.
Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.
Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.
How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
Who is we?
> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely
Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.
Yet
As for consumer-grade solutions, look for the Fujitsu SV600.
I have done it at home for my books since the mid-00s.
But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.
IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.
This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.
The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
Because there’d be much less content created in any media to capture in the first place.
An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.
they should pay the marginal value that the next buyer would buy.
Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?
Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.
Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.
[1]: https://en.wikipedia.org/wiki/Decoupage , since I imagine a number of people won't know what that is.
they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.
having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.
> paying extortion fees to the copyright-mongers
Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s
Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.
And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.
A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.
That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.
Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.
Leave it up for anyone to download and then compensate the copyright holders later.
In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.
For whom? is the relevant question
Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.
On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.
On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...
That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?
Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."
https://www.ala.org/tools/challengesupport/selectionpolicyto...
Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.
What similarities do you see here?
Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?
If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.
Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]
It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.
[1]: https://www.wsj.com/business/media/google-search-publishers-...
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.
A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.
You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
What guarantee do we have that the book contents will be served unfiltered and unaltered?
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
The segment that talks about rare books:
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
I wish...
This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?
Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.
This is highly disturbing news; is this standard practice? What did Google Books do before?
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
https://skybrian-links.exe.xyz/post/1026
Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.
> It is equivalent to book burning in the past. A form of thought control
> > It is equivalent to book burning in the past. A form of thought control
That's only bullshit if you trust AI companies to serve the book contents without alteration.
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
But that would help competitors with training data, which I assume is why they don’t do this.
Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.
If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.
Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.
Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!
There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.
For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
What makes you think they will? What would be the incentives for these companies to do so?
1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.
2. Availability via Google Books or similar.
3. Availability via AI model reference.
4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.
I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.
I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.
What if the perception of those books change over time and are considered masterpieces later on?
Moby Dick was out of print when Melville died 1891 and not a huge success until it got a revival in the 1920s
That’s the whole point because the book shredding is already declared legal.
https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act