Hacker News new | ask | show | jobs
by est31 1 day ago
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

5 comments

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

> Our “rights to read” are diminished significantly with digital works as compared to printed works.

I see it as opposite. I can hand someone the complete contents of a public library on a thumb drive. Delivering that to their door is going to be much trickier.

Reading and distribution (pirate like) has been made MUCH easier, with non-physical distribution. I'm always carrying a 1" thick book in my pocket, that I can read wherever I am. I was ecstatic when I switched over to digital. I could read anywhere!!!

That’s fine as long as you acknowledge that you are breaking the law and infringing the rights of all of the authors of said works who do not want their works disseminated in that way.

If you are willing to give a middle finger to the law, sure, the world is your oyster.

I was speaking to access and ease of sharing, enabling our ability to read, that digital has allowed, not my personal habits. Piracy was not the point of my comment. There are tools to remove DRM from digital purchases, and there are legit stores with DRM free files. And, there's always the option of buying a hard copy then downloading a scanned version (or doing it yourself).
Every innovation since the microprocessor isn't worth saving in the grand scheme of things.

When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

I'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century.
the mass-market-ification of: clothing(world clothelessness used to be a severe problem, paper clothes existed for a reason), AC, light vehicles, solar: all things after the 20th (and even 21st century) that are worth saving. Even with the seeming doom, things do get better.
Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.
Why is the shredding a result of the fair use stuff? I actually don't understand
They’re not legally allowed to keep a physical and digital copy at one time because pf copyright law, and fair use doesn’t cover it as an exception.
> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

Same with the magna carta and the American constitution.

ChatGPT know them, i'd count that as digitalised why keep the originals?

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Because there’d be much less content created in any media to capture in the first place.

Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.
It's kinda amusing copyright always seems to attract discussions about authors rights. It was never about author or their rights. It was and is a set of rules that allow money to be made by publishing books and songs. Authors have a role in that, but so do publishers.

Turns out most authors and publishers suck at their respective jobs. Most books produced by authors are things nobody wants to read. If a publisher's job is defined to be finding works the public likes, they suck at it too. Instead they just publish lots of stuff, mostly at a loss. They make their money from the occasional hit. But only way they can make money from it is if they have a monopoly over publishing it for a while - which is exactly what copyright gives them. Looked at in another way, this "publish lots of stuff and see what sticks" is the way our society discovers what new works are popular. And copyright funds it.

If you take the view that copyright is paying for the discovery and publishing of new works, then consumers paying monopoly prices for 70 years or more is a bad idea. By far the majority of works are commercially dead 1 year after being published. The publishers tend to make their money from the remaining 1%. They might last five years. Very, very few last 20. I don't see how copyright lasting beyond 20 years can be justified with anything other than: "because of the good work he did 20 years ago, he deserves to be paid for sitting on his arse for the rest of his life". In reality, he didn't sit on his arse - he donated to a few politicians - but the outcome is almost the same.

Less of some kinds, more of others. It has been getting easier to create creative works over time, lowering the threshold to make things even for nothing. The netto result would be a bit of an experiment even now, but for sure there are flourishing communities, open source, open content, fanfics, short stories, etc etc. Some hollywood blockbusters even started out as modern internet culture and/or online stories. So there's cross-pollination between the different kinds of communities too.
> It has been getting easier to create creative works over time

When it comes to writing, paper and pens, and then typewriters, have been rather cheap for a long time. Computers made it even cheaper, but it was already so cheap that wasn't the bottleneck. Or take playing instruments, singing even.

But you are totally right about distribution being essentially free now, and with all my qualms about slop and algorithms etc. it's also true that some things become widely known simply because they're really good, sometimes leading someone from rags to riches so to speak, and I can't hate that. I'm sure it would do the same for novels or other "long texts" if more people were into reading those digitally. So while I'd say TV and other stuff killed reading by being louder and more shiny and offering instant rewards, I can't really blame the internet as such for that.

It's the law, logic doesn't enter into it.
> Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

It's faster and cheaper to do all sort of shitty things and people still don't generally do them. AI training may be a numbers game, but the life of none of the involved persons solely consists of AI training. They are persons, not AI training. Likewise, companies only want to make money, fine, but none of the involved people are a company.

It's just shitty people doing shitty things out of sheer greed. This isn't poor people cutting corners so they might survive, bleh to pleading for sympathy for those who are last in queue to warrant it.

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Or physical media, for that matter.

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?

> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

One doesn’t need to pulp the pages after scanning though. After scanning, they could be rebound and put into a library.
That would be a clear case of copyright infringement under current law. You can’t make a copy of a book and then give the original to someone else.
If it's copyright is expired why not?
The books in question are not antique. The Twitter post makes this claim but so far as I can tell it’s not based in fact.

404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright).

”very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases”

https://archive.ph/9MQrK

This is the necessary context distilled down to be concise. Thank you.
Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.
The only issue I see is, how could you tell which are those books?
You can't, without spending a fortune to investigate the scarcity of some of the titles. I do a little work in this space and I don't know of anything destructively scanned that has zero other copies, but definitely some with single digit known copies, where no previous scan exists. Now there is one less physical copy, but a digital copy that none of us can access (except by tricking an LLM).

It's not just the big players buying up these archives, either. There are a lot of smaller players, especially in the OCR space, who are buying up huge swathes of works in languages which have much smaller digital footprints, e.g. Arabic.

We manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is?
The same techniques roughly apply. You can look at all the book marketplaces and see how many of each title are available and it'll give you a rough estimate. If you look across all the marketplaces, all the library catalogs, plus eBay etc and zero copies surface, even going into Worthpoint to look at the last several years of eBay sales, then you know you have a problem. That's how I normally estimate it.

After that you are diving into forum posts etc to see if you can find anyone who has even mentioned owning a copy or having seen a copy.

I don't know what happens to some works. Supposedly thousands, tens of thousands, or sometimes apparently a million or more copies published and yet not a single copy surfaces for years.

It's pretty expensive for the wildlife.

Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.

And if you instituted this, the commenters of this very website would surely decry it as a prime example of government overstep and waste.
Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.
> Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.

1. Why are you here?

2. What is the purpose of this comment?

Books that are shredded can’t be scanned by competitors.
Yeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors.