A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
In many territories like Ireland, Authors are compensated for their inclusion in lending libraries. It tends to be of the pitiful 'music rights organisation' style mechanical reproduction royalties, but it does exist.
The AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors.
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
No, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.
They can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books