Hacker News new | ask | show | jobs
by ACCount37 1 day ago
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

6 comments

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

> Our “rights to read” are diminished significantly with digital works as compared to printed works.

I see it as opposite. I can hand someone the complete contents of a public library on a thumb drive. Delivering that to their door is going to be much trickier.

Reading and distribution (pirate like) has been made MUCH easier, with non-physical distribution. I'm always carrying a 1" thick book in my pocket, that I can read wherever I am. I was ecstatic when I switched over to digital. I could read anywhere!!!

That’s fine as long as you acknowledge that you are breaking the law and infringing the rights of all of the authors of said works who do not want their works disseminated in that way.

If you are willing to give a middle finger to the law, sure, the world is your oyster.

I was speaking to access and ease of sharing, enabling our ability to read, that digital has allowed, not my personal habits. Piracy was not the point of my comment. There are tools to remove DRM from digital purchases, and there are legit stores with DRM free files. And, there's always the option of buying a hard copy then downloading a scanned version (or doing it yourself).
Every innovation since the microprocessor isn't worth saving in the grand scheme of things.

When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

I'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century.
the mass-market-ification of: clothing(world clothelessness used to be a severe problem, paper clothes existed for a reason), AC, light vehicles, solar: all things after the 20th (and even 21st century) that are worth saving. Even with the seeming doom, things do get better.
Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.
Why is the shredding a result of the fair use stuff? I actually don't understand
They’re not legally allowed to keep a physical and digital copy at one time because pf copyright law, and fair use doesn’t cover it as an exception.
> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

Same with the magna carta and the American constitution.

ChatGPT know them, i'd count that as digitalised why keep the originals?

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Because there’d be much less content created in any media to capture in the first place.

Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.
It's kinda amusing copyright always seems to attract discussions about authors rights. It was never about author or their rights. It was and is a set of rules that allow money to be made by publishing books and songs. Authors have a role in that, but so do publishers.

Turns out most authors and publishers suck at their respective jobs. Most books produced by authors are things nobody wants to read. If a publisher's job is defined to be finding works the public likes, they suck at it too. Instead they just publish lots of stuff, mostly at a loss. They make their money from the occasional hit. But only way they can make money from it is if they have a monopoly over publishing it for a while - which is exactly what copyright gives them. Looked at in another way, this "publish lots of stuff and see what sticks" is the way our society discovers what new works are popular. And copyright funds it.

If you take the view that copyright is paying for the discovery and publishing of new works, then consumers paying monopoly prices for 70 years or more is a bad idea. By far the majority of works are commercially dead 1 year after being published. The publishers tend to make their money from the remaining 1%. They might last five years. Very, very few last 20. I don't see how copyright lasting beyond 20 years can be justified with anything other than: "because of the good work he did 20 years ago, he deserves to be paid for sitting on his arse for the rest of his life". In reality, he didn't sit on his arse - he donated to a few politicians - but the outcome is almost the same.

Less of some kinds, more of others. It has been getting easier to create creative works over time, lowering the threshold to make things even for nothing. The netto result would be a bit of an experiment even now, but for sure there are flourishing communities, open source, open content, fanfics, short stories, etc etc. Some hollywood blockbusters even started out as modern internet culture and/or online stories. So there's cross-pollination between the different kinds of communities too.
> It has been getting easier to create creative works over time

When it comes to writing, paper and pens, and then typewriters, have been rather cheap for a long time. Computers made it even cheaper, but it was already so cheap that wasn't the bottleneck. Or take playing instruments, singing even.

But you are totally right about distribution being essentially free now, and with all my qualms about slop and algorithms etc. it's also true that some things become widely known simply because they're really good, sometimes leading someone from rags to riches so to speak, and I can't hate that. I'm sure it would do the same for novels or other "long texts" if more people were into reading those digitally. So while I'd say TV and other stuff killed reading by being louder and more shiny and offering instant rewards, I can't really blame the internet as such for that.

It's the law, logic doesn't enter into it.
> Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

It's faster and cheaper to do all sort of shitty things and people still don't generally do them. AI training may be a numbers game, but the life of none of the involved persons solely consists of AI training. They are persons, not AI training. Likewise, companies only want to make money, fine, but none of the involved people are a company.

It's just shitty people doing shitty things out of sheer greed. This isn't poor people cutting corners so they might survive, bleh to pleading for sympathy for those who are last in queue to warrant it.

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Or physical media, for that matter.