Hacker News new | ask | show | jobs
by kmeisthax 1 day ago
You don't need to shred books to scan them. They make book scanners that will "rip" a fully-bound book no-problem, and even correct for the curvature of the page and binding to give you an equivalent image. In fact, the Internet Archive specifically built nondestructive book scanners[0] for exactly the purpose of which AI companies are now shredding books. The smart / savvy thing to do would be to buy those machines off IA and use them to read the books they're interested in.

The reason why AI companies don't do this is that they're cheap and desperate for training tokens. Same reason why they have scrapers that will happily overload web interfaces for Git repos following links to everything, even though you can just Git clone the repo with far less stress on the host. The AI people are ultimately there just to pillage as much knowledge as they can as fast as possible. Their scraping practices are slap-dash garbage.

[0] https://ones-and-zeroes.ghost.io/scanning-all-the-books-the-...

1 comments

Again, the entire point I'm making is that the choice is dictated by copyright law. They can't legally scan books non-destructively. They can legally format-shift them, i.e. scan them destructively. So this is what they're doing.

Your comment seems to also be regurgitating common misconceptions (to put it charitably) about AI and web scraping.

There is no tenet of copyright law that requires format-shifting be destructive. The term "format-shifting" almost always refers to a non-destructive process; i.e. when you "rip" a CD you are getting a 1:1 copy of the music, but the original CD still exists. If you were legally expected to destroy the disc after ripping, RIAA v. Diamond would have gone a different way and MP3 players would have been illegal.

For AI training specifically, the only standing caselaw is the Anthropic lawsuit. And in that lawsuit, the only thing that was actually in the wrong was maintaining a library of pirated books. That was deemed illegal and Anthropic was ordered to delete those files. But, notably, the judge explicitly said that scanning books to train AI on them was legal, and imposed no requirement to destroy scanned books. I'm pretty sure Anthropic wouldn't even need to retain the physical copies - though there's no caselaw on that in particular, so don't cite me.