Hacker News new | ask | show | jobs
by bayindirh 32 days ago
From my understanding, distilling the model with another model is not illegal per se. Also, the output of the LLM is public domain by law, too.

So, why all this "effort" to protect the model? This is a free market, and moving fast and breaking things is the norm.

If they are so adamant on protecting their IP, maybe they can start by respecting others' IP, so we can start talking about ethics, equality and playing fair.

6 comments

> distilling the model with another model is not illegal per se.

Just because it is legal, that doesn't mean Anthropic wouldn't reasonably want to prevent that from happening (which, from my understanding, isn't illegal either).

I love the asymmetry. When small fish tries to protect itself, big fish hits small fish with "It's not illegal" pole.

When small fish points out that what the big fish is crying about is "not illegal", big fish has the right to be above the law to prevent the problem themselves.

Having values requires equality. They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.

"It's not illegal" is only an argument against lawsuits / law enforcement involvement. Those PoW anti-AI things people put on pages aren't illegal either.
No. From my interactions, I have understood that some people use the same argument to wash their consciences from any guilt. What they do is unethical, but not illegal, and they hide under the same argument to drown the ethical angle.

In other words, being honest to oneself is important.

Anti-scraping measures people utilize are neither unethical nor illegal. That’s the difference.

I’m still frequently shocked by the entitlement people feel to other people’s work/ideas/data/bandwidth/server load, to feed a multi-trillion dollar industry. I find the totally cynical “well when you’re making an omelet…” types to be a bit pathetic, but I understand their motivation— they’re simply greedy. But I just can’t understand the genuine indignation about people attempting to limit or stop ingestion of their own work, even if it’s just for the bandwidth costs. Go ingest your own shit.
I often wonder what those people are like IRL. I'd surmise they're the people that are easy to hate. Greedy and intolerable yet want to be the focus.
It's good to agree that some don't have a conscience, and maintaining an appearance matters more. And appearances change based on what's legal or not.

They could detect the other AI labs and also silently burn the tokens at a faster rate providing fewer tokens for money, which does sound illegal to me.

The comments only further prove that without more regulation around this, big AI wouldn't have a "don't be evil" attitude going forward.

Anyone who called for regulations/guardrails of any kind were shouted down as Luddites who hate progress. We all knew this was going to be a mess but $$$ so screw it right?
> I love the asymmetry.

Much as I hate to defend companies climbing to success and pulling up the ladder afterwards, this asymmetry you note is kind of the whole point a company would want to grow big. Growing an organization has some super-linear costs and generally sucks for most individuals living through it - including the management - but it's still considered worth it, precisely because big entities can do things small entities cannot, and escape the threats from smaller competitors.

It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

> They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.

Yup. Except what they're reaping is insane cashflow and ability to pull stunts like these. We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success, and now are reaping the right to be hypocritical and continue to get away with it.

I've actually heard it quite a few times from different people who want to climb the greasy pole to get heard or resources. Idk it just seems rather soulless and slightly psycho to me. It also seems like that kind of system is rather broken and unstable, if the only way you get impact is to climb up the ladder and whatever that entails.

Change this from humans to companies and I still think it feels slightly wrong.

I disagree with your assessment that large organisations are beneficial.

We can see with our current crop of large organisations that they really struggle to create anything new; most of their new products or services were developed by a small organisation and then acquired. A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

Large organisations are inefficient (everyone has stories of people in large organisations literally doing nothing all day). They are horrible to work for because of the politics. They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

My personal opinion is that we would be much, much, better off if we had fewer large organisations and more smaller organisations.

They are beneficial for those with an equity stake. That much is clear.
Needs a little more precision: not "those with an equity stake": those with a disproportionally large equity stake.

Otherwise it's just an opener for the old excuse of "they might be ruining your life, but it's all good, it's also your pension fund, little man, that's profiting from your life getting ruined, you should celebrate them!"

Agreed. And for oligarchs.
And I disagree with yours :).

Large organizations are necessary if you want things like airplanes and rockets and computers and MRI machines to exist. And if you feel you benefit from those things yourself, then large organizations that create and operate[0] them are beneficial to you, too.

> A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time). Importantly, small orgs and especially startups are very much complicit in this - the venture capital business model in software settled around a symbiosis, where startups create toys (er, MVPs) and growth-hack the shit out of them, in hopes of winning an acquisition or IPO lottery (aka. "exit"), where a big org buys the whole thing for ${a lot}, and enshittifies it further in an attempt of extracting a positive multiple of ${a lot} from the market. Both sides know what they're doing, exits are planned from day 1, and at no point in this process "creating useful products" is ever a driving goal.

Note this symbiosis: it's a recurring theme.

> Large organisations are inefficient

In some ways. Small organizations are inefficient in others. More at the end.

> (everyone has stories of people in large organisations literally doing nothing all day).

Some (not all) cases of this are about maintaining slack in the system, which is necessary for efficiency. A system at 100% capacity is extremely fragile to breaking completely due to tiny, random workload spikes. Breakage is inefficient. Some degree of idle capacity improves overall efficiency.

> They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.

(Of course I'm using a biased sample; I don't know many CEOs of big orgs.)

Symbiosis angle: for abusive practices they can't get away with on their own, big organizations are more than happy to outsource to small orgs and then look the other way.

--

Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.

And yes, you need big orgs to create things like commercial airplanes and MRIs, simply because the big org is a boundary layer, within which you have a non-market based incentive structure, and this lets you build things the free market just cannot reach on its own.

--

[0] - Airports and hospitals are themselves large organizations.

> That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time

I disagree, it's not "modern" economics, it's one half of the Malthusian trap as it manifests in all economics; the other half is that profits tend to zero, both halves are a loss of systemic slack, to reuse the good point you make later.

> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally.

I think this is more like predator/prey size dynamics. One way to keep safe from predators is to be too big to hunt. The regulators and governments are more like predators than their peer-competitors are, cf. "too big to fail".

I actually agree with most of this. A few points:

SpaceX built great rockets before it became large (though this is relative - a small rocket company is a large hairdresser, for example). There is a certain scale required for some types of business, agreed. But getting larger doesn't necessarily make them better.

> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.

Flat disagree with this. Small org CEOs are close to their customers and employees and if they behave like dicks then they get punished quickly. Obviously some still do, because people, but it's harder for a small company CEO to continue being a dick.

> Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.

There is a key step change when the first pure-management layer forms in an organisation. This is the management layer that only have other managers reporting to them, and only report to other managers. So no direct contact with front-line staff or shareholders. Personally, the presence of this layer is what classifies an organisation as "large". It's when the politics takes over from performance as the priority and the organisation starts to lose the connection between what the c-suite want and what the front-line actually do.

And all commercial airplanes, MRIs, anything, were built first by small organisations, and only later by large orgs. Large orgs just can't invent new things unless they form specialist small orgs to do it (skunkworks, or Palo Alto, or similar). Large orgs just don't work like that.

At risk of losing the metaphor, they reaped stuff across all the lands, even ones that were not theirs, and it is questionbale that they even did most of the sowing in the first place
> Much as I hate to defend companies climbing to success and pulling up the ladder afterwards

Based on your post, you don't sound like you hate it at all.

> It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

It's also quite natural to want it to stop at individual human life instead of us getting absorbed by some next bigger thing.

Which I'm fairly sure is also the desire (as far as they can be said to have any) of these animals of various sizes you speak of.

> We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success

Just because they pulled a mirage over people's eyes doesn't mean it suddenly became the "real" world.

Is China the little fish here?
No, the ordinary netizen, who runs their personal web servers, who are hit by crawlers, their content ripped from their hands and their servers chocked during the process.
What Anthropic is doing is illegal in many jurisdictions. I don't know about the legal situation for the Chinese domains they mark, but steganographic data extraction without user consent would definitely be illegal in the EU, for example.
It's not illegal to distil the traces, but it is also not illegal for them to try to stop it.
Imagine an electricity generating company saying that they don't allow their electricity to be used to cold start a competitor's generator.
Do you think software should be regulated as a utility?
AI probably should be. The bulk of its efficacy comes from the work of “everyone else” (in loose terms). AI also aims/hope/threatens to replace such a large number and range of jobs that it probabky should be a commons.
Wise words. And probably in a few years more people will think the same, but now most are blinded by the gold rush or the hate for it.
> Do you think software should be regulated as a utility?

I do think that AI models that were trained using the biggest intellectual property heist in human history should be a utility for all, yes.

I would 100% support this as long as the same is true for human works. Every artist is trained on thousands of years of art tradition. Copyright wasn’t a thing for the first many millennia of artistic creation, and we still got Bach and Michelangelo.
> Michelangelo

Did you know that Michelangelo was the first "label" for inventions and art? He sourced his material from other people and published it for them.

So it's kinda ironic (maybe on purpose?) that you mentioned him.

> as long as the same is true for human works.

and it is.

Anyone can study Michelangelo or Bach, and learn their style.

Depends on the software.

In this case, the companies that make and provide AI models that are increasingly used to interact with me on critical things (banks, public sector services) then yes.

Abso-fucking-lutely they should be regulated like crazy.

In fact I'm really surprised by the amount of people that are not worried by how many parts of their lives are being handed over to be managed by a probabilistic system that is controlled by a private company with next to zero oversight.

There must be a greater liability than "oops, you're right to push back"

Software, no. But maybe eventually AI and tokens are a public utility.
AI companies say people will buy intelligence like they buy electricity.
> Also, the output of the LLM is public domain by law

Why so? Also there is a lot of code in ironically claude and ChatGPT that’s generated by LLM . Yet I haven’t seen the public domain code

The code is not eligible for copyright. If they do not give you a copy of the source code, that does not matter. And if you don't know which parts were generated by LLM, you can't safely reuse the code.
> And if you don't know which parts were generated by LLM, you can't safely reuse the code.

I speculate this could be a real issue in future copyright infringement lawsuits.

The plaintiff bears the burden of proving that the code they claim is copyrighted by them actually is copyright. If it is known that large parts of it were generated by LLM, they’d need evidence to demonstrate sufficient human input to establish copyrightability. If they’ve kept highly detailed traces of the development process, that could be rather straightforward; if they haven’t, it could be really difficult.

Now, that’s true in the US, which never accepted mere “sweat of the brow” as a basis for copyright; the UK courts have, and most of the Anglosphere follows the UK on this more than the US.

The other factor: when dealing with an (almost) trillion dollar corporation, even if you’ll win the legal argument, they may bankrupt you with legal fees before the argument is ever properly heard.

But I suspect the precedents on this topic are going to be established by lawsuits involving far smaller actors.

(IANAL and I speculate only for myself, not any present, past or future employers.)

> The code is not eligible for copyright.

This is very much not what the linked case established.

According to the link:

"The US Copyright Office and federal courts require human authorship for copyright protection; works created solely by AI are not eligible for registration under the current rules."

The Supreme Court declined to consider a challenge to this rule, and so for the moment at least, the rule remains in place.

This means that companies leaning heavily into their LLM use may very well find that they do not, under the law at least, actually own any of their code. As I've read elsewhere there's every possibility that AI code will be the asbestos of the Software Engineering world. Something we'll be trying to get rid of for decades, once everyone comes to their senses.

I think the word "solely" is going to be a tunnel you can drive freight trains through.

So with the asbestos analogy, we encase the fibers in resin and call the whole thing copyrighted.

Or in other words: there's a big difference between public domain and copyleft and it looks like whoever came up with the asbestos analogy was underestimating that difference.
Please explain. That is exactly what the linked case established.
It established you can't assign copyright to the LLM itself. That's very different.
> If they are so adamant on protecting their IP,

What they are trying to protect doesn't qualify as intellectual property. Only 4 categories of IP exist: (1) copyrights; (2) patents; (3) trade secrets; (4) trademarks.

The capabilities embedded in model outputs don't qualify. Machine-generated outputs are ineligible for copyright. They aren't covered by patents. They aren't trade secrets, because the model companies are selling them rather than keeping them secret. And of course, trademarks are conceptually inapplicable.

This leaves the model companies with contract law (ToS) which is pretty inept because it can't bind third parties. And technical measures, like the ones being discussed in the article. And, of course, politics.

Frankly, I think it's pretty ridiculous to even think that models can be protected from being learned from. I feel the Stanford Alpaca team demolished that idea 3 years ago.

The hypocrisy of the pro-AI mega corp arguments makes my head spin. For three years they’ve been using the example of a human reading books and then outputting creative works influenced by them as analogy for training AI on copyrighted works. Now suddenly we’re supposed to not draw the same parallel about a hypothetical person who learned from Claude and is now outputting creative work based on it.
> So, why all this "effort" to protect the model?

Because it's their model and business and they are free to use the free market to do exactly that?

That's their free market rights too. If you don't like it, use another model (which they would be fine with).

> Because it's their model and business and they are free to use the free market to do exactly that?

I mean, nothing stops distillers to find better ways to distill, either. Meaningless cat & mouse games.

> If you don't like it, use another model (which they would be fine with).

Thanks, I use none. It's peaceful this way.

The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not, which is what China has been doing. Countless thousands of separate requests abusing the service (which is not a simple static HTML feed, but an AI service request) for every kind of query to soak up the results.

The post is about what's in the local code, but for a long time there has already been modifications made to the request outputs from the major cloud services as they work together to both curb adversarial distillation and to degrade the quality of training China can get from that distillation.

It's likely not to make the answer wrong or bad, but to make it so that any model trained on the output would not gain the benefit of the model's reasoning generalization skills as easily and also identifying markers that might even link back to request IDs.

The techniques talked about in this post are naive and simplistic, largely because they are released publicly.

It's not as much about protecting IP as much as it is about slowing China down or being able to track the effects of abuse. So many people are talking about greedy company this, greedy company that. The world is not made up of caricatured giant money pigs wearing suits with monocles and gold watches. That is a children's view of Marx's exaggeration on free markets. Bad, greedy people do exist, but if that is your only hammer for every nail then you have a problem.

The upper-bound for how good these models can be is so crazy that it is essentially dual-use military applicable to an extent most other technologies are not. It's not only cyber attacks or biological weapons. Most people are not even built to understand the possible threats.

Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.

> Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.

Reminds me of this comic: https://xcancel.com/tomgauld/status/571994690289061888?lang=...

None of the superpowers in this world is innocent, and like MAD, more countries have the capability, the better.

I know some of the things CCP do/did. I know some of the things US does/did. I'm from neither, so I don't take sides.

AI's use has been confirmed, or more precisely boasted by two countries in two different wars, and China was not one of these countries.

We have seen the effects of "if they don't know them, they can't exploit them" mindset of NSA for years. Keeping information/technology private is neither beneficial, nor possible. It's only a temporary moat-ish gap. Not a definitive solution.

Certainly the world is full of actions and reactions, nothing is happening in a vacuum. You don't have to be from a country to take sides, but presumably you have some kind of moral compass, some kind of values around personal freedom or the worth of a human life.

There can be a very real cost, because one side comes from an ideology with a history that wants to conquer the entire Earth which caused World War 2 while the other side is trying to prune the planet like a bonsai to prevent it from descending into total chaos to preserve some sense of international order.

Europe was constantly at war, and we helped stabilize it. Middle East as been constantly at war, and if Iran can be sorted then it will be the closest to some sense of peace it's been in a long time.

We used to be in Japan, Philippines, Germany, Vietnam, South Korea, Iraq, Afghanistan and so on. How many are US territories? None. We aren't out there to conquer the globe and take land. We're usually fighting other people's wars for them, because they're up against better resourced opponents. Meanwhile China is over there building artificial islands, ramming other country's ships, creating ideological police stations in countries around the world to harass people and engaging in the most widespread international interference campaigns in human history.

They do not treat their people well and they do not have free speech. The internet is flooded with their propaganda now, because they have a human numbers advantage.

It's true that given time most advantages are temporary, but there's always that slim chance we could slow them down until the CCP collapses and they could become a more normal country.

You sound like you've swallowed pro-US-propaganda hook, line and sinker.

The reason the middle-east is at constant war is because colonialist machinations. Same goes for south-Saharan Africa. And the US is a big colonialist player, just ask Vietnam, South-America, Iran, Afghanistan, etc. They all have been attacked by the US because of US colonial interests. If anything, one could make the argument that the PRC is treading much more lightly than the US.

That said - I'm not defending the PRC by any way; it's a state-capitalist hell hole that's suppressing workers by denying them any ability to organize and whose political class is purely focused on furthering their own interests and that of the moneyed elite, the common person be damned.

The thing is - so is the US.

If you think those were about colonization, you would be well served to examine history closer.
Ah yes. The exceptionalism argument.

There's the good "us" and the bad "them."

I didn't make an exceptionalist argument, but if any country's behavior and values can be measured compared to others, you will always be able to make some kind of decision about where those fall in terms of goodness or badness. Do you not believe in good and bad?
I believe in good and bad.

I don't believe in US good, non-US bad. I also don't believe the same about my religion or my political party, for that matter.

How you measure depends on weights you assign (cultural system of values) and what information you use (media bias).

You can rank in the extremes (e.g. North Korea as worse than Belgium), since they come out that way by almost any set of information and values. Comparing the US to most other countries, there isn't a clear ordering. If you believe the things you wrote, I think the other comment summed it up well: "You sound like you've swallowed pro-US-propaganda hook, line and sinker."

Most countries have similar propaganda, by the way.

> The CCP is darkside material.

And which country has "Black Sites" peppered around the world to detain and interrogate people they don't like?

The US doesn't have black sites anymore and when it did, the interrogation techniques were chosen to avoid physical harm. The results were bad, we didn't like it here in the US even if they were extreme measures for extreme times and so we shut it down. It had a high error rate and generally didn't reflect what we thought was right.

Meanwhile the CCP regularly abducts its own citizens and executes more people than the entire world combined.

> The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not

This is literally what the "training AI on copyrighted works is just like a human learning/getting inspired" crowd has been arguing though.

Literally. People have been literally saying that it was wrong because they did this "learning" at scale in a dishonest way.

In some ways it's an offshoot of the honest benefit of search engines already crawling all this content. That has its own conflicts, like just how much of a page's content should you reproduce in the results before it's basically considered stealing their content without benefiting the site itself.

There is a balance to strike, both in search engine fair use cases and AI fair use cases. The major cloud LLMs do double as web search engines now, though they didn't originally. In many cases there's no reason left to click the links they sourced from.

That is a legitimate concern. At least within the US, I think there are nuances around fair use and contract law. A lot of companies are getting paid for having their content used in these models, but many websites had no particular rules you had to abide by and the content was simply public. I think if you're operating under an agreement, then even if there is fair use or public domain content being reproduced by the site you are still bound by that agreement.

Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used. These AI services are obviously very different, but there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage.

I'm not exactly comfortable with the mass scale that everything was soaked up to train these models even within the umbrella of search services, but I also admit that a lot of the usage was probably quite legal. The potential displacement caused by the resulting trained models on artists or writers is almost its own facet. In practice, whether they ONLY trained on strictly legally acquired fair use content with no errors and paid agreements to acquire even more content than they already do or not, there was enough legally accessible information for fair use that there was no escaping some kind of impact on artists, writers, etc.

With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

Well it was also problematic when the search engines started quoting the websites in such a way to disincentivize people from visiting the actual website.

> At least within the US, I think there are nuances around fair use and contract law.

The concept of "fair use" as it exists in the US-law system is completely dysfunctional (see e.g. nearly every educational music channel on YouTube), so utterly biased to favour large corporations, that there's very little room for whatever "nuances" you believe exist.

> Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used.

Yes 300 year old paintings are public domain. Indeed there are certain rules for the people/institutions who digitize them. It's not "they have some say", there's actually nothing mysterious about it and it is not similar to Anthropic's copyright heist at all because nearly all of the books they copied were not more than 100 years old.

> there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage

well where I live, there are laws about what a "service" can claim to "lay out as acceptable usage" instead of the other way around ...

> I also admit that a lot of the usage was probably quite legal

Let's disagree on that. I think it wasn't a lot and the vast majority was not legal. How do you think the LLMs "learned" to speak all these non-English languages? Unless your point is that it's probably quite legal to treat foreign IP like that. Which it may very well be in the US, especially if the corporation is large enough, but imvho it's still wrong.

> With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

And with any bad luck, these AI corporations will hold frontier models hostage for the rest of time.

I honestly don't want to put that up to "luck".