Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
61% Positive
Analyzed from 4548 words in the discussion.
Trending Topics
#books#book#copyright#copies#law#copy#amazon#fair#training#library
Discussion Sentiment
Analyzed from 4548 words in the discussion.
Trending Topics
Discussion (77 Comments)Read Original on HackerNews
Before my lifetime, copyright law all but choked and killed the public domain. And now everything is stale and the same.
This particular battle was lost with Google Books and the attempt to make the world's largest library. The modern library of Alexandria.
But then copyright lawyers got involved to get their pound of flesh. And here we are.
I am upset about the destruction of knowledge. Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago. I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them.
Everyone is impressively short sighted.
Most paper books don’t make it a thousand years, let alone a hundred. Your local library is probably purging piles of books every year and nobody sheds a tear.
> I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them
It would be cheaper to buy a second copy of the book and put it in a different warehouse where it can continue to go unread as it did before the AI companies bought it. The cost of non-destructive scanning and rebinding books is crazy high.
The basic point of a book is for people to (a) read it and (b) synthesise it into a comprehensive world model. And people can burn their own books if they want to, they own the thing. The only possible complaint here seems to be the scale and it seems like a big challenge to say that is a problem given that knowledge from books is allowed to be used at scale.
AI content is not exactly helping everything feeling like Extruded Product.
I'm more than willing to bet that in 5-10 years they'll wish they didn't destroy the books, specifically the rare ones.
To be clear, I'm not trying to condone their behavior. I hate it. But I'm trying to show that even if you had their same ethics it's still dumb. I hope workers are secretly stashing the books away. If you're one of the people in charge of destroying the books, you have a cultural duty to preserve them. I'm willing to bet people will go to great lengths to help you do it secretly so you can continue to keep them safe and keep your job. If you're at Amazon, or any company where this is happening, you have a duty to the world to make efforts to preserve the books. Steal the PDFs and lock them away. Create backups in your company. Tell your bosses you don't think this is right. Don't sit silently while this happens. Silence unfortunately is enabling. Unfortunately silence isn't a passive action
This has been less true every year since Usenet and Geocities. Much of the stuff on their was pirated, but that was the "stale and same" stuff you're complaining about anyway. The rest of the stuff was all kinds of original, good and bad. There's never been more total "content" and more total variety. You just have to look around for it (but that's also never been easier).
Letting AI companies try to profit off of all that creativity forever while choking the ability of the creators to make future revenue off of it is exactly what would actually lead to a "everything is stale and the same" situation. Preserving variety and novelty of new creation would look like putting in new restrictions to reduce this not-envisioned-when-the-laws-were-made sort of read-once-slop-forever usage.
That does not make any sense! Really might be high time to scrap it all.
On the other hand, absolutely yes we should abolish all IP law.
Build a lending library, or build a bonfire, or do anything else with them that you choose.
They're your books. Have at it.
Copyright ways. Form change probably should be compensated. Even if it means that you wouldn't be able to format change your own media. Or that such activities wouldn't be allowed without compensation over certain threshold.
And in times of substantial inequality, they lag yet further.
The only relevant case law for AI training in the US is the rulings in the Anthropic lawsuit presided over by Judge Alsup. That lawsuit ruled that it's infringement to build a shadow library from pirated books; but NOT to train AI on those pirated books. The only point where destructive book scanning even comes into play is that Anthropic also had a book scanning program alongside their piracy, Judge Alsup said that program was not infringing, and Anthropic happened to be destroying books. At no point did Alsup say that leaving the books whole would have infringed copyright - it was never even considered as it was outside the scope of the lawsuit.
Now, if Anthropic were to non-destructively scan books, store them in a library, and sell the books on, that could be infringing. All the case law about format shifting presumes the owner retains the original. So Anthropic would likely have to hold onto books, at least the ones they wanted to train on, until they were done training on that book[0]. But they do not have to destroy them permanently. They are destroying these books specifically because it is cheaper to do so than to use, say, the Internet Archive's own custom-built nondestructive scanners.
[0] I am absolutely furious about how much this sounds like "fair use is just an extra license you get when you buy a book", and I would much rather have had Judge Alsup just say AI training is not fair use instead.
> The copies used to convert purchased print library copies into digital library copies were justified, too, though for a different fair use. The first factor strongly favors this result, and the third favors it, too. The fourth is neutral. Only the second slightly disfavors it. On balance, as the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.
I do not think the ruling would have gone this way if the books were not destroyed, as many would assert that Anthropic would retain them to sell later, otherwise. The fact that destruction is mentioned so pervasively in the decision suggests it is an important fact to consider.
https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
Are you referring to the 2023 ruling? If so, that was filed in 2020 explicitly because the CDL was not limiting the total number of copies in circulation:
https://storage.courtlistener.com/recap/gov.uscourts.nysd.53...
“We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility” (404media.co)
164 points | 3 days ago | 321 comments
There's also a fair use consideration that favors noncommercial use compared to commercial use, but that's not the only question, and the statute doesn't say clearly how to combine the fair use factors.
But it's possible under the Copyright Act that some commercial uses of copyrighted works could be considered fair uses while some noncommercial uses could simultaneously not be considered fair uses.
The whole "copyright infringement is theft" meme seems to have stuck.
20 years ago it was all "information wants to be free" and "infringement is not theft; being deprived of speculative profit is not a loss."
How the turn tables.
He's lead Metallica through the "Copy our tapes! Let everyone hear our music!" early days, through the "They're stealing our master recordings!" Napster and Senate hearings era, and come all the way back 'round to "I’m just happy that fucking anybody cares about what we’re doing and shows up to see us play and still stream or buy or steal our records or whatever.” [https://consequence.net/2023/09/metallica-lars-ulrich-stream...]
2 out of 3 ain't bad, I guess.
"Information wants to be free as long as I benefit from not paying."
I "just" compared prior experience and current experience specifically.
this is a complicated way of counting how many of each word is in the book.
clearly fair use.
does the author think cutting up a book is a copyright concern? theyre buying the books, its up to them what to do with their copy. if you want the books preserved, maybe fund your libraries to get a copy or two?
Not the case here, but it's also important to keep in mind that laws too, can be unjust and immoral. Much of humanity's progress has involved the abolishment of immoral laws.
This is the whole problem. A lot of people assuming they can speak unilaterally and with authority on what other people value or should find valuable.
It's disgraceful. What happens to the old hacker ethos? I'd never thought I'd see the day when, on a site called hacker news the majority of posters are defending the right of a large corporation to destroy public goods permanently just to be able to charge people more for "intelligence as a service".
So you are saying to abolish copyright?
Who's to say? I'll check back in 10 years :)
At any rate, I ought to pick it up again and finish reading it...
Lastly, there's a misguided belief among some that it's somehow less of a copyright violation if the source is destroyed, but it's a violation either way unless the entity doing the scanning has permission from the copyright holder to make the copy.
Then setup a fund to compensate copyright holders.
I guess publishers do have a right to pull certain works from print though.
For a refresher: The Internet Archive scanned books then shared them online, trying to claim that converting them into digital copies qualified as a derivative work. You don’t have to be a lawyer to see how that claim doesn’t hold up to any scrutiny.
The LLM companies are not redistributing the works. They are using them for training. This blog post calls it IP theft, but it has actually been litigated in court already. Using books for training does not qualify as theft or redistribution, even though some people have different opinions about the moral angles.
One of those same lawsuits also extracted a huge settlement from Anthropic for using digital downloads from pirate sites. The conclusion was that the only acceptable way for them to use the books is to buy them and scan them. The courts forced it to be this way.
The current hand-wringing about the destruction of books is based on claims that they’re doing it to rare books that are also valuable. So far nobody has been able to actually provide an example of one of these books that is supposedly ultra-rare, but also valuable, but also only available at one of these book sellers that sells these things in bulk. We’re supposed to assume that one of these books might actually be super valuable but also super rare and also only available at these places they’re buying from.
For the record, i don't much have issue with llms and such.
But... the wholesale destruction of printed books is a bridge too far.
There’s no giant “used” section at Barnes and Noble. Libraries only have so much space and they have to cycle through books as the years go by. This is easily the best fate these books could have received.
Google Books tried. They lost. If you want this fixed, talk to your congressman, not Amazon.
Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?
Your granpa once wrote an obscure book about this amazing way he'd found to cure diabetes.
Corporation A buys all existing copies of the book, scans them, destroys the originals, and sets up a commercial business offering a monthly subscription to alleviate diabetes pains with this new method they claim they discovered.
Person B borrowed the book from a municipal library, Xerox'd it, and keeps a free ledger, open to all who want to read old books, as a way to safeguard free access to the world's knowledge.
Do you think there could exist any possible logic by which some people would defend person B and try to stop Corporation A?
This is not happening. They want the text, nothing more. Depriving others isn't the goal, based on all empirical evidence about their behavior. It's about one each of every ISBN.
They're intended to be analogous of reality.
Which real-world company is said to be buying all existing copies of a book, here?
However, the recent headlines have been about "rare" books, of which it can therefore be assumed that few copies remain, and fewer of those (possibly none) would be accessible to the public if they are in the hands of collectors or long-time buyers.
As an author, I know some of my own books didn't sell and probably exist by now in small enough quantities that if you destroyed the boxes in my apartment, and bought the last single copy still available used on Amazon, nobody's ever getting their hands on the other few that might be surviving in a couple of basements here and there.
So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.
I'm sure people would prefer if Amazon also offered library-like access to the book, but how is it theft? I'm not asking about "any possible logic" - he called it theft.
Ok, but it's a bit different because they are not selling "the book" they are selling a fundamentally different thing for which the book is consumable input. Let's consider a batch of 5000 books. For kicks let's assume the known copies are N=1 for those books. Today, though only a few people can do this at a time, a person can purchase one of those books once read exactly the text in that book, and keep doing so to their heart's content. Or maybe a library buys it and now a rotating legion of people can do that for free. Now let's say amazon buys and scans them, destroying them in the process and refusing to release scans to avoid giving competitors a training edge. Now:
- You can never get the information as written again. At best you'll get an LLM output approximation/mutation of it. If this was the expression of a real human beings lived experience, that's kind of sad and goes against one of the spirited aspects of human existence, to leave a legacy, doesn't it? - Amazon is going to charge you per token every time you want to access that information. - Price demand for the information is now tied up with general demand for LLMs rather than the actual book, either reducing or greatly inflating the cost to you in addition to the now recurring charges.
So, now you (a) can't actually ever access that book as it was written (b) need to pay continually to access an approximation of its contents (c) and possibly more than the book is worth since now it's "value" in a price sense has been absorbed into general llm inference costs.
Not to mention you may have eradicated the last extant copy of a person's memoirs, but I guess it's pretty clear people in this industry don't care at this point. I hope one day your entire life story is ground up and consumed in some data farming operation and you are all summarily forgotten.
"AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so.
Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying.
Or they don't care about any of that and just wanted to highlight the hypocrisy.
[0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest.
https://news.ycombinator.com/item?id=49310725
https://news.ycombinator.com/item?id=49068738
It does though. A lot of people seem to think "it's a book!!!" means it takes on some mystical intrinsic value. That's bullshit. There's an absolute mountain of worthless and near-worthless books. I bet the vast vast majority of these books are those books - after all that's what Amazon wants. A ton of human written text as easily as possible. They obviously aren't buying first editions of Oliver Twist, or even second editions of Harry Potter.
> We’re not revealing the titles of the books included in the shipment we tracked
Yeah... because then it would reveal how unimportant they are. This sort of outrage inflation is counter-productive. People see through it, lose trust, and then when something bad actually happens they won't believe you.
https://en.wikipedia.org/wiki/Outrage_industrial_complex
And come on. They're not revealing the titles because that would reveal which bookseller they worked with.
I've read a lot of unknown old literature that has little to recommend it as far as literary history goes, but it all has a certain charm, and at the end of the day it was all the work of a human being who walked this earth just like we did and captured their thought on the page.
What if their family has an interest in the preservation of that text? What if they simply don't yet know of it, or couldn't afford it?
Yes all hypotheticals, but I think a lot of people are underestimating the potentially permanent damage being done to the public good just because the actions are not illegal.
Everyone should be uncomfortable with a single company unilaterally pillaging the public good for its own ends. Period.
IDK about you geniuses on hacker news, but I'd much prefer to pay a few hundred dollars for a rare book once to be able to read it than pay amazon token costs monthly to get the LLM's chopped and screwed regurgitation, without any recourse to even knowing where the information comes from.
Individually you might be right. Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
This is cultural violence on a massive scale.
I just searched again now and found 3 used copies on AbeBooks for about US$11 incl. shipping. Absolutely made my day.
I don't know if adult me will enjoy it beyond the nostalgia, but there's a coconut cake recipe in there that I will never forget.
Like you said, just about nobody on planet earth would care if the billionaires ingested and then destroyed those final copies, but I'd have been gutted if they'd beat me to it.
quantity has a quality of its own
>Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
ask your aunt for the recipe, look it up online with a search engine, ask an LLM by describing how it tastes and the ingredients you remember, this is really not that complicated
>This is cultural violence on a massive scale.
could you be any more dramatic? you're acting like we are actually losing anything when in reality we're gaining from this, you'll just be able to ask the AI for the recipe instead of wasting hours and hours hunting down some old book and ordering it online for an exorbitant amount
Your example is very flawed because if the LLM gave you the exact recipe then it would be infringement, but many recipes need to be exact. LLMs aggregate information so wanting a very specific recipe (to the point of hunting down an old cookbook) is a non-starter because it the LLM is really only guessing.
The whole problem here is:
- Not everyone agrees on the utility / value of AI. - A small handful of companies control frontier models and get to shape how these inputs are ultimately put to use. - As Sam Altman put it, they want to charge you, in perpetuity, for "intelligence" as a service. Contrarily, anyone can access a book usually for a one time purchase fee and read it forever. You'll be paying per token to look up auntie's recipe every time instead of just, you know, reading the book.
Not to mention you are erasing the voices of human beings who had lives and may have related their experiences and emotions in those texts. That anyone would be comfortable with a corporation they know nothing about destroying potentially unique information of the lives and stories of real human beings, locking it away forever as one granular component of a stochastic content generator, the human expression in its actual form never to be accessed again, is absolutely appalling to me.
I think you should reflect more on who will really benefit, who has the control, and why humanity as a collective is potentially losing as collateral, without having much say about it.
Do you even hear yourself?
These behaviours would be described as obsessive and pathological in any clinical setting, but when they're done for profit it's somehow normal.
Just like China does not have communism. They have too much free-market activity and private ownership for it to classify as such.
If you look at China and think "That's not communism", just be aware that they may also look at the west and think "That's not capitalism." Both are right. In fact both systems have converged.
China took the better aspects of communism and capitalism. We took some of the worst aspects of both... Except maybe free speech; but people are working hard on removing that one!
It makes sense; China came out of an era of extreme government control which has been reduced over time (compared to what it was before). In the west, government control has been on the way up and we can't even imagine how bad it can get.
Wage labor is then the perfect structure to ensure that most resistance from the bottom is suppressed before it can even really start. When you depend on the corporation to succeed for your subsistence (since there is no adequate social safety net to fall back on) taking a principled stance and refusing to destroy great works of literature becomes a highly costly endeavor for most people.