Hacker News new | ask | show | jobs
by mapontosevenths 25 days ago
If I read Harry Potter I will remember some parts verbatim. Others I will tecall in only an abridged and lossy way.

Does that make my brain copyright infringement? Does Disney now own all my output forever because some small part of me now has Harry Potter embedded?

6 comments

Can you remember every part? Can you do this for every book in a library? Can you remember all that forever?

If you just ignore anything that's inconvenient for your argument, you can make any argument you want.

>Can you remember every part? Can you do this for every book in a library? Can you remember all that forever?

None of those are relevant factors when it comes to copyright law. You don't get a pass for copyright infringement just because you're not copying the entire work. Same goes for a copy that's transient. You can't set up a bootleg movie theater in your home, even if you delete the movie file afterwards, and there's no trace of the movie aside from the viewers' vague memories.

> None of those are relevant factors when it comes to copyright law.

And yet they very much are. US copyright law has the concept of "fair use" in 17 U.S. Code § 107 [0]. I'll paste here for your benefit, #3 is the one I referenced as most obvious but #1 and #4 are also very relevant:

  (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
  (2) the nature of the copyrighted work;
  (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
  (4) the effect of the use upon the potential market for or value of the copyrighted work.
Naturally remembering some parts of a legally purchased book verbatim is fair use. "Memorizing" the entire library obtained via torrents and incorporating that in a commercial product that can output all that content doesn't sound like fair use to me.

The US justice system is too captured and corrupt at this point to take as reference because decisions there are bought by the highest bidder. But for the purpose of this discussion let's not play dumb for the benefit of trillion dollar corporations.

[0] https://www.law.cornell.edu/uscode/text/17/107

>And yet they very much are. US copyright law has the concept of "fair use" in 17 U.S. Code § 107 [0]. I'll paste here for your benefit, #3 is the one I referenced as most obvious but #1 and #4 are also very relevant:

If you're going to invoke fair use, that opens up a whole can of worms on what counts as transformative. The google books case and the google thumbnails case shows that you can make near verbatim copies of works at scale and still be considered fair use.

>The US justice system is too captured and corrupt at this point to take as reference because decisions there are bought by the highest bidder. But for the purpose of this discussion let's not play dumb for the benefit of trillion dollar corporations.

This is begging the question. The original question is whether ai companies are getting special treatment. You can't then use that as a premise to say that the courts are tilted towards ai companies. Not to mention it's questionable how ai companies were suddenly able to corrupt all the judges, some of which were appointed decades ago, even though they only got rich a couple of years ago.

Look, first you were wrong with the confidence of an LLM and claimed an argument that was literally in the definition of copyright fair use had absolutely no relevance whatsoever for copyright. Even now you are surprised that I invoked fair use on reading a book. That was to respond to someone who "brilliantly" brought up reading Harry Potter [0] as evidence that the law allows any extent of "memorization" and reproduction of copyrighted material.

Then you switched to a barrage of questions on the premise of words in my comments that were neither written nor implied. If you muddy the waters just enough maybe everyone gets lost in there.

> The google books case and the google thumbnails case shows that you can make near verbatim copies of works at scale and still be considered fair use

Now maybe we agree "reading Harry Potter and remembering some lines" is indeed fair use, but you decided my argument is still not relevant to create a distinction between "reading a book" and "feeding it all into an LLM" because of an vaguely related exception. For better or worse thumbnails are a copyright violation according to some courts [1]. But looking at the big "Books" decision (this is the one you meant?), did you check out the court's opinion [2]? Why would you believe the two cases are substantially similar? Just because they're both big tech? Just for yourself, from the definition of fair use and referencing that opinion, do you see any significant differences between "Google Books" and "big LLM"?

> You can't then use that as a premise to say that the courts are tilted towards ai companies

The highest bidder is what I said.

> Not to mention it's questionable how ai companies were suddenly able to corrupt all the judges, some of which were appointed decades ago, even though they only got rich a couple of years ago.

You're getting creative" about what I wrote. "AI companies"? They are just the big corrupting agent of the day, and nobody with deep enough pockets had "revolutionized" the legal areas they're working in to this degree until now. Tech in general has been doing it for a few decades already. Other incredibly powerful industries have been doing that in their respective areas for even longer. "Suddenly"? The US justice system has worked exactly like this for so many decades when it came to the interest of very deep pockets. "All judges"? I said "the system" because all judges don't have the ultimate power to ultimately decide on things.

I'm surprised at your surprise that reading a book is fair use, and that courts have been "captured" and beholden to economic interests above justice for so long we forgot when it started.

[0] https://news.ycombinator.com/item?id=48774664

[1] http://www.linksandlaw.com/news-update59-thumbnails-germany-...

[2] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,...

> Can you remember every part?

No, and neither do LLM's. They're trained on vast quantities of data and retain only a fraction of it.

You might think of it as very, very lossy compression that generates new outputs rather than the original input unless something unintentional happens.

> If you just ignore anything that's inconvenient for your argument, you can make any argument you want.

I'm not. I just understand how it actually works. You either don't understand or are deliberately ignoring that what you just said is literally and technically untrue to make some sort of political statement.

Does the law really not distinguish between mechanical processing of data, and humans learning from it? It seems surprising to be if every person who read a textbook is copyright infringing. It also seems surprising if something like a lossy compression algorithm is enough to protect you from copyright law.

Somewhere between the two a line must be drawn… where we’d want to put that line, I guess, if up for quibbling. But it doesn’t seem obvious to me.

>Does the law really not distinguish between mechanical processing of data, and humans learning from it? It seems surprising to be if every person who read a textbook is copyright infringing. It also seems surprising if something like a lossy compression algorithm is enough to protect you from copyright law.

The google books and google thumbnails cases have so far upheld that even mechanical reproductions are allowed, depending on the context/usage.

To me the distinction hinges on the output being transformative enough to be considered a new work. I think that most of the time LLM output is.

Sometimes they go a bit wonky and overtrain on specific phrases which can result in verbatim copies of brief sections of coontent. Thats a bug, not a feature.

If you write out the parts or recite them for other people to hear, yes it's copyright infringement.

Humans reading or watching copyrighted material isn't considered "making a copy" for the purposes of copyright law. Machines doing so generally is.

Further, why has my brain's searing remake of Snow White as a gritty murder mystery gone unscathed by Disney lawyers? Surely their negligence has diluted the Snow White trademark!
This analogy is disingenuous because by comparing the human brain to the machine, it ignores _scale_. Scale is absolutely important in copyright law. As a matter of fact, copyright law is among the various profound impacts of the---wait for it---printing press, a _machine_ for the mass production of books.
So if I watch a LOT of Disney movies THEN they own my own unique output forever?
yes it is if you write it down from memory and sell it. Exactly what LLM companies do