Hacker News new | ask | show | jobs
by bonoboTP 10 hours ago
If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.

You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.

5 comments

The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.

The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).

Copyright protects against reprinting or reproducing the wording and expression, not using the idea expressed in there in novel contexts.
AI reproduces copyrighted work exactly in many cases, so it clearly infringes copyright in this sense

The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely

Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI

Are you a lawyer that has tested this in court, or is this just what you want the reality to be?

As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.

I mean, its theoretically possible that a court might rule that if an LLM outputs an exact or lightly modified piece of copyrighted work, that it won't be copyright encumbered. We'll end up in a situation where copyright doesn't exist anymore, because you can always claim that its been laundered through an AI. This seems terribly unlikely to me, because 1:1 transformations (eg copying an image into memory) are already established to count as making a copy for legal purposes, there's strong precedent around piracy

There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true

The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph.

Derivative work or transformative? It's not the same.

AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad

AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case

You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here

As long as it is not substantially similar to training set it should be OK to reuse ideas. Ideas should not be protected by copyright or we can't create anything. We should not accept "vibe copyrights".

Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.

Are you sure you didn’t have RAG enabled and it wasn’t putting the paper in its context for your prompt? I find it hard to believe that anything but the most commonly published papers/code would exist directly in model weights.
Suddenly it's "copyright infringement" to count the amount of times one word occurs after another word. I find this whole thing so amusing.
AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement
> The output of it is also a derivative work,

A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.

Can you point me to the statute that makes it ok for humans to learn from copyrighted work, but forbids AI? I’ll wait.
This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.

The issue is not who owns knowledge, it's how it benefits humanity.

They can't apply for a patent in those categories, so they have no way of preventing others from reading their text (or using a program to process the text, and compute statistical properties), and using the ideas in other contexts (without re-expressing the same text).

If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.

The world would be quite different if AI companies had to create the knowledge they trained on, rather than consume that knowledge freely given away. They're not known for freely giving away their produce either, I don't know why you think the ire should be pointing in this direction.
I like open models for sure. I support free software as well. But using published knowledge to solve new problems was never disallowed, even for profit. Today, a for-profit company, e.g. a gigantic Big Pharma company can have their employees read chemistry and biology papers and use the knowledge gained from it to improve their products and processes and make more profit without paying a cent to the authors (beyond what they may get - likely nothing - due to the potential paywall).
Patents shouldn’t exist.
Patents were invented to incentivize disclosure of inventions and prevent extended secrecy. In exchange for the public disclosure (which allows others to experiment with the idea without selling yet), you get to keep exclusivity for N years.
And?

Currently patents are predominantly used to prevent interoperability and impose costs.

“If you don’t lock your bike, you can’t prevent others from taking it for a ride. You enjoy the mobility attached to the idea you’re a bike rider who rides a bike but then you play this game when that bike would actually be useful as opposed to sitting in the bike rack all day”.
Nope, and I'd also download cars.