RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
67% Positive
Analyzed from 2187 words in the discussion.
Trending Topics
#books#book#information#don#more#old#labs#scanning#interesting#years

Discussion (72 Comments)Read Original on HackerNews
I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.
But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.
Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.
We were taught respect for books, but not everything is worth preserving.
People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.
Yeah, now instead old books being turned into toilet paper, we'll get turned into toilet paper.
And Sam Altman will become richer than God, and isn't that what really matters?
But don't worry! You'll still have access to ChatGPT until your savings run out.
I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are
I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.
Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.
The goal here is to have all human knowledge in a single file, which is pretty neat IMO.
I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.
They're scanning millions of books.
It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.
There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.
I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.
Books are more likely to be about a specific topic or story or time or setting and be more information dense
LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.
They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).
OCR Scanned for training, then tossed away or burnt. Great for nature.
1) The labs don't want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.
2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that's why the labs are starting to purchase by the pallet, it's because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.
3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don't get destroyed or thrown away.
Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don't expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.
> the books being burned by these misanthropic lunatics are not available in any other medium
prove it, name one title
> this is not about fascination with some particular mediumn of transmission
it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.
> But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.
https://www.offthebeatenshelf.com/blog/pulp-fiction-is-real
BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.