The problem of handling recurring tasks predictably comes down to the probabilistic nature of LLMs, which are based on next-token prediction.
I founded a company called Aide where our goal was to help support teams reliably deploy customer-facing agents without worrying about poor interactions. The first problem we needed to solve was making them deterministic and eliminate the variance that comes naturally with base models.
Getting them to always adhere to brand policy, eliminate hallucination, and stay grounded in data was a fun challenge. Proud to say that we’ve devised a solution that runs well and it’s worked out quite nicely in compliance-heavy and regulated environments.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today)
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.
The internet made this even worse by celebrating people who only have to say the right thing.
The increase in people who believe that perception is reality has the logical consequences of people shouting their opinions loudly enough so that it becomes 'the right thing'
Knowing and enacting are two distinct things, "you should do what I want you to do" means you're already doing the wrong thing, and yes, there is an agreement on what the right thing is, its what is ethical, that is, what non profits and archivists are forced to do.
I can’t believe that in the multipolar world of 2026, where hundreds of conflicting world views coexist, where universalism is being disproven on a daily basis, it’s still possible to read things like “there is agreement on what the right thing is”.
This is as blatantly false as claiming that the Earth is flat, and the fact that there is no such agreement (descriptive moral relativism) has been firmly established in philosophy for well over a century.
If you think ethical arguments like "murder is bad" or "human knowledge should be preserved" and a demonstrable falsity like "the earth is flat" are equivalent arguments then you are so far gone you might as well be a flat earther.
Blind futurists and AI cheerleaders scare me on how cavalier and how many crimes against humanity they ignore.
> The example book of Old books of agriculture is probably not that important today
If I may be flippant, not to you but to the sentiment, skill issue.
We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.
And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era?
What if... you just want to learn what your ancestor was doing at such and such time?
What if you want to find clever techniques for robot arms to work with food crops in space?
Or, heck just the alpha from a hedge fund point of view of finding old climate patterns and... :)
Your ability to make the most of knowledge is only limited by your imagination.
> Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
I don't understand what you're trying to say here.
> The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
"No-kill shelters" are very rare and very controversial (thanks to PETA); most still have to kill by law (overpopulation), and they instead tend to focus on not taking in more than they can handle. They are very controversial because they are more of a "moral" than a "practical" solution to animal overpopulation. They try their best to let animals have their "normal natural life to the end". They are just more focused on the animals' well-being, kind of like a library for books, but they still have to kill if they overpopulate.
The main article is written like it is the "end of the world", just like your random rambling about "climate instability" is just the natural course Earth has been on for its existence, with melting ice.
Random article about powder and stone, completely irrelevant.
The old book mentioned was just a farmer's book; you can find all this history in the area's historical records.
"Skill issue." How old are you? I think it would help you more if you contacted a shelter and talked about what they really do.
Better sources than Wikipedia that explain what they do; it is, of course, different from shelter to shelter, but euthanising due to overpopulation is still part of it. (Like too many books in a library, they need to make room for other books that people want, if it just end up roting on the shelf and nobody loans it over years, either destruction or storage, most don't have space for storage)
Because destroying the book makes it possibly inaccessible permanently because we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that.
If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up. Hopefully someone will at least be able to purchase their digital scan and hopefully the scan is of decent quality and clearly indicates the provenance of the text.
> we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that.
Books are considered super high quality training data. Anthropic has no reason to get rid of this data that 1. They’ve spent a ton of money on and 2. Will remain useful indefinitely for training LLMs.
> If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up.
The same applies for a digital scan? The information isn’t any more likely to be lost.
Yesterday you could have purchased one of these books and read it. Good luck doing so today.
I don't know why you're so quick to assume this will all just "work out" such that the scans are ultimately accessible. Arguably the most likely scenarios are either Anthropic survives and holds them away in perpetuity or Anthropic fails and they are sold off to the highest bidder who does the same.
I don't think that changes the message at all. If a librarian doesn't like routine book destruction, they would still feel worse about this different thing that is happening.
Weeding (deselection works) is a fundamental part of collections management. Every trained librarian is going to understand this.
The interlibrary loan system has mechanisms in place to make sure the member libraries keep two copies of each work in each region. Collection managers consult these databases during weeding to make sure they don't deaccession the last copy.
If the monograph was never collected by a library and it gets caught up in a destructive scanning project then I guess it was pretty "rare" in a literal sense. "Rare Book" in library land is sort of a term of art and I'm not sure if the books in these destructive scanning projects meet the criteria.
I'm not a bsky person so I didn't click though but library books aren't "retired" to a farm upstate. My SO works at a library, they are mostly shredded. This is a nothingburger
Just go to a local library and ask a librarian or anyone who has a lot of books and tries to give them away; sadly, in most cases, they pick the valuable ones, and the rest just get sent for destruction (Burning).
The difference everyone is missing, is that now there is commercial incentive to burn books aka destroy them after scanning. It's now a profitable business to do so, not something that only has to be done to clear up space
Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
> Correct. I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades. It would have eventually found its way to the discount rack and then the dumpster. At least this way we're getting some money in that we can use to buy actual cool books/collections, pay salaries and bills, etc.
That data is useful as a source of scientific knowledge even if it's not current. Although it's probably already online, they probably don't want to download it and risk getting another copyright lawsuit
I think under first sale doctrine you have a much stronger case with destructive scanning. Google Books, HathiTrust, and Internet Archive's book scanning project have had a lot of legal expenses.
I didn't fully understand why but apparently there is also a legal reason to destroy the books, it makes it them less likely to be considered copyright infringement.
Not really. Selling the book onward does seem legally dubious but legally nothing (yet) prevents you from storing the book in a warehouse. obviously it’s cheaper to dispose of them.
> This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.
Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.
If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.
Are you under the impression they scan the book, train on it, then destroy the digital copy? Because that's not what's happening. They scan it, and hold it forever to train future models on. The scan still exists, not available to the general public but that's no different than if they had bought the books and kept them in a private library closed to the public.
>it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge
This sentence reminded me of “The Things” by Peter Watts. The Thing in that short story believes it’s actually the good guy and decides to commit “violent integration” for the sake of humanity.
I’m not sure I buy the links premise that Anthropic is doing this with malice, but if it were, would be an eerie parallel.
They're scans. Those are human-readable. They probably won't make them available to the public, which is the exact same state they would be in if they bought the books and just put them on a bookshelf in a warehouse.
A lot of them are rotting. We are not talking "one of the three living copies of the first edition of Joyce's Ulysses". Rather "1956 statistics of the cultive of yuca in 'some small village from Mexico': a boring analysis". Those books have value to train LLMs as they are 100% free of AI text, but has been collecting dust (or rotting) in someone's room for decades, and no human is buying them even for 10 cents.
Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
They are ridiculous because the reality of the situation is boring. Boring doesn’t drive engagement. So, all the clickbait headlines and ragebait comments imply scandal. If they didn’t, they wouldn’t get attention.
After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.
This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.
Does china developing ai powered missiles sound ridiculous? It sounds ridiculous to me because the US is surely doing the same, It would be boring if we knew the us was doing the same. It is emotional to think about china destroying the US when it is just the antithesis of American exceptionalism. It reminds me of what Dario said about their stance on open source.
That doesn't sound ridiculous at all (maybe reductive but not ridiculous). And the US doing the same doesn't make it any less interesting either. In fact the US's use of AI in developing target banks in a recent conflict caused quite a lot of consternation when publicly revealed. Ukraine has also recently used AI enabled autonomous drones to make an entire area one big "kill zone", no humans required. Not so ridiculous or boring if you ask me.
>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
No. Almost none of the books you donate to the library get to the shelves. If they can sell it, they will. Otherwise they are thrown away.
I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".
As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices.
Our local library-connected biannual book sale puts a large dumpster, about the size of an 18-wheel truck trailer, outside the warehouse where the sale is conducted. At the end of the sale it gets filled to the brim with waste inked cellulose, and sits there open in the weather until it's taken away for disposal.
No, it's a demonstration of the lack of value of all that waste. On the last day of the sale (which decreases prices steadily over the several weeks it lasts) you can take a paper grocery store bag, fill it with books, and pay just $1. And yet enough books in the sale, enough to fill a large dumpster, don't even make that cut.
I think it hits different when you're buying large quantities of used books with an eye towards ones which aren't readily available online, and with the specific intention of destroying them.
If all these companies were doing was burning star wars tie-in novels and harry potter sequels nobody would care. That's not their goal because they already have those in their training set. The whole point here is to find rare or underappreciated books from the pre-digital era which nobody ever made publicly available in a digital format.
BTW destroying them isn't even necessary for scanning. It's the easiest way because removing the binding and turning it into a flat stack of papers solves many problems but there are actually dedicated book scanners designed to hold open the book while its photographed, and un-curling pages in post-processing was already a solved problem long before people were using AI to correct images.
Realistically the vast majority of those books haven't been touched for half a century and won't ever be read by a human again. It is a price to destroy these, but it's a price I would pay for progressing models towards AGI and the billions of lives it will save.
> When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.
Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.
As someone who has gone to many many used book sales over decades… many of the books at a sale never get sold, guess what they usually get tossed in the dump. This includes your local library book sales. Books are heavy and worthless. Cheaper to throw away the ones that nobody picks up in a sale.
I know it comes at a shock but truly most books are absolutely worthless.
No, that’s the hyperbolic reaction clickbait wants from you.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.
Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.
And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.
I heard from the first stories that virtually all of these books being ordered have ISBN numbers. Books that are rare that have ISBN numbers are rare because no one wanted them 99.9% of the time. Somebody wants every book, but you'd spend many, many years finding that somebody.
A flagship LLM today is trained on tens of trillions of tokens, the equivalent of hundreds of millions of 100k word books. No human has that kind of appetite.
As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.
Exactly. It’s the funniest thing. These books are only being bought by AI companies. It seems pretty clear they only have value to AI companies. Otherwise all these people bemoaning the loss of rare books would be… buying them.
“How dare these companies buy rare and valuable books that nobody else values enough to buy” is a self-canceling argument.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).
> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Hyperbole much?
Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!
If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
A small change in the copyright law would fix this problem. Something like:
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
Puts the burden on government to store what is probably 90% worthless material.
Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.
There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.
You don't need a massive government program for this.
Just post on r/DataHoarder: "Free 16TB NVMe SSD to anyone who indexes and mirrors the entire out-of-print 20th-century physical archive." The problem would be solved by next Tuesday. With probably 10x redundancy and people willing to do it for free for fun.
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
> Mandatory deposit applies to any work published in the United States. This requirement does not apply to works first published in a foreign country until they are published in the United States. Copyright registration is optional, but it provides additional legal benefits and fulfills the mandatory deposit requirement with the submission of the required copies.
Since we are talking about a US perspective do you have evidence that backs this up? It just comes across as an empty statement. The government is the will of the people and I personally like the idea of fixing copyright instead of making the government store how to use windows 95 books.
Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.
Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?
How does this help anything, except create more work to throw in the trash?
That revolves around print on demand for books that are out of copyright or where the copyright has been abandoned.
>
Background
Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions.
If you want a physical copy of The Sorrows of Satan, you can buy it from them.
> The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment.
> Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment.
That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.
Actually - most jurisdictions that issue a publisher a unique root ISBN number have a stipulation that anything new published using that number must have a copy sent to them.
Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.
Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive should be available for research use (including ML training) at cost.
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
----
> Acceptable Formats for Deposit of Electronic Works
> The deposit of electronic works is arranged with the Acquisitions & Deposits division.
> For electronic-only works, submit the best edition in accordance with the formats listed in the “Electronic-Only Works Published in the United States and Available Only Online” section of the Best Edition Statement (PDF, 135 KB).
> For works subject to a grant of special relief, unless otherwise specified, the Library will accept an appropriate “preferred” format listed on the Library of Congress Recommended Formats Statement. Such files must contain no measures (such as digital rights management [DRM] technologies or encryption) that control access to or prevent use of the digital work.
> For more information about electronic deposit, see the above FAQ “When can I make an electronic deposit of a work?”
---
> When can I make an electronic deposit of a work?
> Works may be deposited in a physical format in accordance with the Best Edition Statement, which can be found in Best Edition of Published Copyrighted Works for the Collections of the Library of Congress (Circular 7B) (PDF, 135 KB).
> Works may be deposited electronically in certain circumstances:
> The Copyright Office issues a written demand for an electronic-only book or serial. If your work is published only online and the Office sends you a written demand for mandatory deposit of the work, you must deposit the work electronically.
> The Copyright Office offers you electronic deposit as an alternative to depositing a physical copy of the work. If you receive a letter offering special relief to deposit a work in an electronic format instead of sending physical copies, follow the instructions in the letter or agreement.
Right, that's why I said we should make copyright require deposit and registration (like it used to). Publish your work without its copyright ID for people to use to reference the LoC database? It is now public domain.
(iii) Automatic protection A key feature of the Berne Convention, and thus also of the TRIPS Agreement, is that copyright protection - unlike most other forms of IPRs - may not be subject to any formality of registration, deposit, or the like. This principle is contained in Article 5(2) of the Berne Convention, that has been incorporated into the TRIPS Agreement.
So? Of all the aggressive things the US forces upon the world (or just unilaterally does, ignoring agreements), undoing its own bad policy would be a drop in the bucket, and would be doing some good for once. I'm sure if we just did it, others would respond tit-for-tat and require registration of our material, and then mission accomplished.
(And at the end of the day, sovereign people are never required to do anything. The concept of international law is an oxymoron)
I don't believe it is a bad policy that copyright on anything that is copyrightable is automatic (I don't need to register this comment with the Library of Congress).
It's not to make AI training easier; AI training is already happening. It's perfectly easy for them.
It's to preserve our heritage and knowledge. Automatic copyright is what causes information to be lost. I don't suppose you're going to submit your comment to LoC or otherwise keep it available for 70 years after you die? Does everyone remember to submit their code they publish?
We're not losing books because of AI companies. We already lose them because the law makes it so only groups like Anna's Archive can save them.
If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.
> it would be easier for them to distribute the work after the copyright of the work expires
Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years.
Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120).
Certainly. But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.
The library of congress does distribute some digitized works that are out of copyright. And it does digitize some works for archival and distribution, but having additional works digitized for (eventual) public use could be nice.
> But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.
It is not at all clear that this is true.
The number of books published every year is growing rapidly. According to Bowker the number of books published in the US every year has increased ~15x in the past two decades [1].
Because of this, I suspect that the median age of the books we are discussing is below 30 years.
In theory the Library of Congress could "lend out" digital copies like some library systems do. This would be especially helpful for rare books since it is less likely multiple people would want the same book concurrently.
I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)
If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.
"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony
Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."
"For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
The third fair use factor favors fair use for the purchased library copies converted from print to digital."
Why would a company keep the hard-copy around at the risk of it being inadvertently given away, resold, etc.? It's a huge outstanding liability given that the illegal copying of works -- the other part of that case -- is what they settled out of court for some huge amount of money. Destruction is the only thing that makes sense.
I'm old enough to have been around when DCMA legislation was under discussion. Many people were dead-set against it and raised concerns over matters exactly like this. In Rainbows End (2006), Vernor Vinge wrote about a similar scenario where a robot went through the university library shredding books, and scanned the shredded pieces to recombined them into a digital archive.
Anthropic may be doing shady things and may have even done this on their own recognizance, we just don't know. As it stand, this is 100% a consequence of US copyright law, much of which was written by large corporations to protect their own assets.
I agree fully that it makes logistical sense. But it is not a legal requirement, and they should not be permitted to use that as an excuse to wash their hands of their own decisions.
I think I agree with you in spirit. I don’t like what these companies are doing, and Anthropic’s actions can for the most part stand on their own. Copyright law is just a special interest of mine, and I do think it’s important to recognize what external incentives exist and what they prioritize. Because other companies will act in similar manners under the same incentive structure, and the problem is going to cascade and magnify if it hasn’t already. There are active court rulings setting precedence for this behavior - take note!
The ruling was that what they did was within the law, not required by law in every detail. They cannot resell the copies. I saw nothing in that ruling that required their destruction, because it is not required.
A person can digitize their own books without destroying the original. So can Anthropic. They are choosing to destroy the books for easier scanning and trying to palm off the blame for it.
The ruling was that training is fair use but still they can't make unauthorised copies to enable training. They had to pay a huge fine for each unauthorised copy. So now they aren't making unauthorised copies.
1 point by jonhohle 0 minutes ago | edit | delete [–]
You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.
That’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.
They cut the bindings off and feed loose leaf books through a document scanner. It would be difficult to store and probably unsellable, it's probably going in the trash.
To legally retain these scans, you must own the original book. You can sell or give away the original book (sans binding) but it's legally dubious as to whether the scan can be transferred along with it. So if Anthropic has no interest in storing thousands of loose leaf books, they are likely destroying both the original and scan as soon as possible.
At the end of the day, the only thing of value Anthropic has is the trained model which is definitely transferable.
> I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity
Look, others talked about how this fetishising of paper books is quite silly (though I don't like it when the destructive scanning is just so one could feed it into a chatbot) but I have to say, all I can do after reading the above sentence is laughing bitterly. Anthropic is an American for-profit company, any talk about "working towards the benefit of humanity" is just marketing lies, and it's always incredible to see people treat those seriously.
Depends on the point being made about “historical precedent” and the lessons to be drawn from such.
Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.
It's a fair point. I honed in on "the burning of" (original comment) versus more generally thinking in terms of "the loss of", because parent context here is "AI companies destroy…".
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
I wouldn't even compare AI to any of one. Rather, AI will even contrast with some of these, democracy for example, if we stop critically thinking and delegate everything to AI (= big corporations that control it) it cannot end well to me. Same thing if we continue building datacenters that generate tons of CO2 for doing stuff we can do with well, our brain (that is the most energy efficient computer on earth!) or even better not do at all (because we don't need AI slop), air conditioning will not be enough I fear.
Computers and computer programs did work much more reliably when they were written by humans, a modern computer system written mostly with AI has bugs that not even in systems that were in use in the 80s (and some of them are even still in service today!) had.
Do you want a planet where the human being is at the center, or the machine? Do you think machines has to be at your service, that you are the one ordering it what to do (that is programming it), or you think you are to be a servant of an AI that takes all the decision for you? Because if you think the second, you want MATRIX or SKYNET, well I don't want MATRIX or SKYNET to be fair (not that it's even possible btw, since AI has anything intelligent in it beside the name, it's just a glorious copy/paste machine).
The idea that these corporations or any of the literal sociopaths that work for them give the slightest bit of a shit about "benefitting humanity" is hilariously naive. The one and only thing these entities care about is money, and making as much of it as they can. If they could get away with it, they'd commit every crime that exists if it meant they get a quarter of a percentage increase in their quarterly earning reports.
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
They're not destroying rare manuscripts or incunables.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
It's possible for a datacenter to use scarce well water irresponsibly.
They don't have to, and the vast majority don't.
The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.
is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want.
i dont think the CCP cares whether there's data centers in rural america.
Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers:
* sound and infrasound under x DB
* no air quality change
* must pay to build out electrical infrastructure
etc
its not to the CCPs benefit or loss to make sure the data centers are built well if they get built
That AI consumes it at an abnormally high rate, presumably.
The claim never made sense to me either, I can only assume those that regurgitated such claims never worked with HPC or even general datacenters before.
Was recently talking to a (non-technical) friend about this, she was surprised after talking about the "insane water use for AI datacenters" when I responded that open-loop cooling is pretty rare for a datacenter and I've never actually seen it used before, versus closed-loop (or just regular air-based cooling) which has no real noticable water consumption.
Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere.
2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.
That AI data centers are drinking up local ground water for cooling. It isn't (or wasn't) nonsense, though. It was/is a real thing, though it seems to be on the out in favor of closed loop cooling after the massive and still on-going public outcry.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
shameless plug but that's exactly what we're building at Aide (https://aide.app), I would love to chat and see how we can potentially help if this is something you're interesting in learning more about or deploying!
would love to learn more about how you're using Freshdesk and the types of queries you're responding to! I'm building a customer support app that uses openAI's GPT3 to generate automated responses and it's been working great so far!
The holy grail [https://beepb00p.xyz/pkm-search.html#future] of this really resonated with me and fully mirrors what I've been thinking about the past few months. In my observations, it's input capture, information organization, and subsequent retrieval:
Information Capture:
Input Capture - You’re going to have all-encompassing tracking and recording of all activity, but want configurable privacy on the extent to which you want your daily conversations and observations of external things you encounter and are exposed to. Capturing input needs to be holistic and incorporate all properties of encounters and new information.
Potential sources of input:
Vision — point of view recording, see snapchat spectacles, etc as primitive examples.
Audio (voice notes and multi-party conversations) - voice calls, video, etc. and other forms of audio transmission where there is more than a single party in the interaction.
Digital interactions
You will need to keep track of web pages you visit at what times
Conversations you see on Twitter, etc.
Properties and cues must be extrapolated from the information that is captured on input, in the case of audio, transcriptions are sufficient for transcription and retrieval purposes, however since video is a visual medium, it includes significantly more properties that need to be accounted for.
The aim here is to identify sufficient data points (cues) that are subsequently represented in such a way that they are easy to search across things you have encountered but only seem to recall a certain property or cue from. This is because of the fact that human beings tend to remember things in fragments, for instance, you might remember a certain color on a page that you visited within the last 6 months and nothing else.
So long as you are capturing sufficient input and actions then you should be able to go back to any given point in time. How and where are you going to store this information? Storing everything is going to be a large amount of data. The essence of the information and context must be preserved. If you want to wind back to an arbitrary position in time with the original context intact, you want to retain as much as you can in the most efficient manner possible, so determining which data points to retain is essential. (Once the content structure has been figured out, this will be viable).
Examples of Primary Cues:
Time - humans generally keep track of things in a linear time-based fashion.
Color - invokes emotion and is memorable.
Physical Location - the efficiency of information retrieval is highly influenced by the location at which it is originally synthesized, encountered, and stored.
Keywords - the default conventional mode. Can and should be extracted from video/imagery and audio.
Imagery - search for images based on their contents and ambience.
Potential Secondary Cue — Music - see historical associated input and actions while certain music was played.
(What else?)
Meta Cues — Subjects - Automated tagging of keywords/encountered content.
Any combination of these queries is possible, but ultimately the killer feature is the ability to backtrack through time to find a certain piece of information that is made available thanks to the always-on recorded nature of your interactions with the physical and digital worlds combined.
Knowing what to store, and how, + displaying it needs to be worked on further.
http://onemodel.org, described elsewhere here and moreso at that site, tries to model arbitrary knowledge and has a vision encompassing any kind of info one wanted to be tracked (again, more at the site).
(Edit: If you have possible future interest, there is an announcements list.)
You really need to let us try it out without having to sign up. Imagine how many users that think "oh, yet another img -> html/css tool, I wonder how questionably it works" and then get discouraged by having to take a chance on something they have no idea how well/if it works.
Interesting research, in my opinion, this been long coming as the next phase in the evolution of computing and society in general. It's how tech is going change the world to the point of 24/7 always-on connectivity; once the tools, applications, and entire ecosystem is created with more consideration to the human element of the system.
We're actually applying to YC's Summer 2016 batch with a fully automated application deployment process to accommodate for developers, by shifting focus to deployment considerations that have not been eliminated.
I'm based in Saudi Arabia, and considering the deserted climate and environment, there are literally no persistent water sources (aside from a few select wells that are over-exploited by bottled water manufacturers).
I don't know what you mean by "feasible". Saudi Arabia neither has the population (28mil vs 38mil) of California nor the Farming needs (I don't have numbers on Saudi farms but i'll guess its somewhere near zero).
There are a number of desalination plants along the coast of California and a few more scheduled to be built. However they're very expensive to build and to operate and their yield isn't where it should be.
Supplying 50% of Cali's water via desalination is not at all feasible
When people talk about California's water problems they make it sound as if there isn't an easy solution, but there is. The real core of this entire issue is not the methods but more the cost, it is ultimately a conversation about saving money NOT about some finite limitation on water in real terms.
California could solve this issue with a pen stroke, it just might hurt their farmers, which is really what all the concern is about. If water doubled or more in price (which is realistic), that is expensive for farmers who need a ton of the stuff for their crops. So will supermarkets pay 30% more or will they look abroad?
I actually think even with a higher water bill, it will still be cheaper for US retailers to buy US produce. Shipping that stuff by ocean isn't exactly cheap with the price of oil. I think where it would hurt US farms is their exports to Europe in particular, Europe is in a geographical position to buy from either the east or west, both by ocean. So if US/California crops go up in cost they might just buy them from someone else.
But let us not pretend that either shipping water in from other states OR just distilling water isn't an option for California, because it is. It just might hurt farmers and make them less internationally competitive.
This is part of what is needed. Water rights in California are as old as the state and extremely convoluted. Those with the older water rights have a practically guaranteed supply and generally irrigate in remarkably wasteful ways. Breaking the old water rights and increasing the price of water would push farmers to move to Israeli style computer controlled drip irrigation rather than current methods of just spraying tons of water over the field. Gov Brown was talking about bring that technology over and pushing it hard into farming...
Feasible for human drinking but I assume not farming (citation required). Which goes along with the story of "People in cali are not going to die of thirst, but a lot of farmers might go bankrupt"
After having a look and reading your interview with marketingstartups[1], it is quite an interesting idea.
It strikes me as an evolved incarnation of Buzzfeed[2] in that the entirety of its content is user-generated, whereas this wasn't the case with Buzzfeed until recently, and even now, I believe user submissions are subject to moderation.
I don't think you should give up yet. I wouldn't. I would start by simplifying the design (sure, it looks nice, but it feels over-ornamented), defining stricter guidelines for posting (e.g. encourage "Top/Best X" type posts in lieu of general free-form lists, and figure out other ways to start driving traffic to it.
I wish you guys the very best and again, don't give up yet.
PS: Please get in touch if you'd like to talk more.
That's exactly what we're trying to make; a kind of Buzzfeed 2.0, where the "2.0" stands for improved functionality and quality of lists. That's our vision anyway.
I saw your site, looks very cool. Perhaps I could mail you our story?
I founded a company called Aide where our goal was to help support teams reliably deploy customer-facing agents without worrying about poor interactions. The first problem we needed to solve was making them deterministic and eliminate the variance that comes naturally with base models.
Getting them to always adhere to brand policy, eliminate hallucination, and stay grounded in data was a fun challenge. Proud to say that we’ve devised a solution that runs well and it’s worked out quite nicely in compliance-heavy and regulated environments.