Hacker News new | ask | show | jobs
by gleenn 28 days ago
Whether or not you find Anthropic's behavior bad, theybhave been very loudly stating the foreign labs have been distilling their models for a while now. This seems like an obvious response to me that would be a mechanism to make that obvious.
16 comments

From my understanding, distilling the model with another model is not illegal per se. Also, the output of the LLM is public domain by law, too.

So, why all this "effort" to protect the model? This is a free market, and moving fast and breaking things is the norm.

If they are so adamant on protecting their IP, maybe they can start by respecting others' IP, so we can start talking about ethics, equality and playing fair.

> distilling the model with another model is not illegal per se.

Just because it is legal, that doesn't mean Anthropic wouldn't reasonably want to prevent that from happening (which, from my understanding, isn't illegal either).

I love the asymmetry. When small fish tries to protect itself, big fish hits small fish with "It's not illegal" pole.

When small fish points out that what the big fish is crying about is "not illegal", big fish has the right to be above the law to prevent the problem themselves.

Having values requires equality. They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.

"It's not illegal" is only an argument against lawsuits / law enforcement involvement. Those PoW anti-AI things people put on pages aren't illegal either.
No. From my interactions, I have understood that some people use the same argument to wash their consciences from any guilt. What they do is unethical, but not illegal, and they hide under the same argument to drown the ethical angle.

In other words, being honest to oneself is important.

Anti-scraping measures people utilize are neither unethical nor illegal. That’s the difference.

I’m still frequently shocked by the entitlement people feel to other people’s work/ideas/data/bandwidth/server load, to feed a multi-trillion dollar industry. I find the totally cynical “well when you’re making an omelet…” types to be a bit pathetic, but I understand their motivation— they’re simply greedy. But I just can’t understand the genuine indignation about people attempting to limit or stop ingestion of their own work, even if it’s just for the bandwidth costs. Go ingest your own shit.
It's good to agree that some don't have a conscience, and maintaining an appearance matters more. And appearances change based on what's legal or not.

They could detect the other AI labs and also silently burn the tokens at a faster rate providing fewer tokens for money, which does sound illegal to me.

The comments only further prove that without more regulation around this, big AI wouldn't have a "don't be evil" attitude going forward.

> I love the asymmetry.

Much as I hate to defend companies climbing to success and pulling up the ladder afterwards, this asymmetry you note is kind of the whole point a company would want to grow big. Growing an organization has some super-linear costs and generally sucks for most individuals living through it - including the management - but it's still considered worth it, precisely because big entities can do things small entities cannot, and escape the threats from smaller competitors.

It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

> They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.

Yup. Except what they're reaping is insane cashflow and ability to pull stunts like these. We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success, and now are reaping the right to be hypocritical and continue to get away with it.

I've actually heard it quite a few times from different people who want to climb the greasy pole to get heard or resources. Idk it just seems rather soulless and slightly psycho to me. It also seems like that kind of system is rather broken and unstable, if the only way you get impact is to climb up the ladder and whatever that entails.

Change this from humans to companies and I still think it feels slightly wrong.

I disagree with your assessment that large organisations are beneficial.

We can see with our current crop of large organisations that they really struggle to create anything new; most of their new products or services were developed by a small organisation and then acquired. A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

Large organisations are inefficient (everyone has stories of people in large organisations literally doing nothing all day). They are horrible to work for because of the politics. They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

My personal opinion is that we would be much, much, better off if we had fewer large organisations and more smaller organisations.

They are beneficial for those with an equity stake. That much is clear.
And I disagree with yours :).

Large organizations are necessary if you want things like airplanes and rockets and computers and MRI machines to exist. And if you feel you benefit from those things yourself, then large organizations that create and operate[0] them are beneficial to you, too.

> A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time). Importantly, small orgs and especially startups are very much complicit in this - the venture capital business model in software settled around a symbiosis, where startups create toys (er, MVPs) and growth-hack the shit out of them, in hopes of winning an acquisition or IPO lottery (aka. "exit"), where a big org buys the whole thing for ${a lot}, and enshittifies it further in an attempt of extracting a positive multiple of ${a lot} from the market. Both sides know what they're doing, exits are planned from day 1, and at no point in this process "creating useful products" is ever a driving goal.

Note this symbiosis: it's a recurring theme.

> Large organisations are inefficient

In some ways. Small organizations are inefficient in others. More at the end.

> (everyone has stories of people in large organisations literally doing nothing all day).

Some (not all) cases of this are about maintaining slack in the system, which is necessary for efficiency. A system at 100% capacity is extremely fragile to breaking completely due to tiny, random workload spikes. Breakage is inefficient. Some degree of idle capacity improves overall efficiency.

> They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.

(Of course I'm using a biased sample; I don't know many CEOs of big orgs.)

Symbiosis angle: for abusive practices they can't get away with on their own, big organizations are more than happy to outsource to small orgs and then look the other way.

--

Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.

And yes, you need big orgs to create things like commercial airplanes and MRIs, simply because the big org is a boundary layer, within which you have a non-market based incentive structure, and this lets you build things the free market just cannot reach on its own.

--

[0] - Airports and hospitals are themselves large organizations.

At risk of losing the metaphor, they reaped stuff across all the lands, even ones that were not theirs, and it is questionbale that they even did most of the sowing in the first place
> Much as I hate to defend companies climbing to success and pulling up the ladder afterwards

Based on your post, you don't sound like you hate it at all.

> It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

It's also quite natural to want it to stop at individual human life instead of us getting absorbed by some next bigger thing.

Which I'm fairly sure is also the desire (as far as they can be said to have any) of these animals of various sizes you speak of.

> We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success

Just because they pulled a mirage over people's eyes doesn't mean it suddenly became the "real" world.

Is China the little fish here?
No, the ordinary netizen, who runs their personal web servers, who are hit by crawlers, their content ripped from their hands and their servers chocked during the process.
What Anthropic is doing is illegal in many jurisdictions. I don't know about the legal situation for the Chinese domains they mark, but steganographic data extraction without user consent would definitely be illegal in the EU, for example.
It's not illegal to distil the traces, but it is also not illegal for them to try to stop it.
Imagine an electricity generating company saying that they don't allow their electricity to be used to cold start a competitor's generator.
Do you think software should be regulated as a utility?
AI probably should be. The bulk of its efficacy comes from the work of “everyone else” (in loose terms). AI also aims/hope/threatens to replace such a large number and range of jobs that it probabky should be a commons.
Wise words. And probably in a few years more people will think the same, but now most are blinded by the gold rush or the hate for it.
> Do you think software should be regulated as a utility?

I do think that AI models that were trained using the biggest intellectual property heist in human history should be a utility for all, yes.

I would 100% support this as long as the same is true for human works. Every artist is trained on thousands of years of art tradition. Copyright wasn’t a thing for the first many millennia of artistic creation, and we still got Bach and Michelangelo.
Depends on the software.

In this case, the companies that make and provide AI models that are increasingly used to interact with me on critical things (banks, public sector services) then yes.

Abso-fucking-lutely they should be regulated like crazy.

In fact I'm really surprised by the amount of people that are not worried by how many parts of their lives are being handed over to be managed by a probabilistic system that is controlled by a private company with next to zero oversight.

There must be a greater liability than "oops, you're right to push back"

Software, no. But maybe eventually AI and tokens are a public utility.
AI companies say people will buy intelligence like they buy electricity.
> Also, the output of the LLM is public domain by law

Why so? Also there is a lot of code in ironically claude and ChatGPT that’s generated by LLM . Yet I haven’t seen the public domain code

The code is not eligible for copyright. If they do not give you a copy of the source code, that does not matter. And if you don't know which parts were generated by LLM, you can't safely reuse the code.
> And if you don't know which parts were generated by LLM, you can't safely reuse the code.

I speculate this could be a real issue in future copyright infringement lawsuits.

The plaintiff bears the burden of proving that the code they claim is copyrighted by them actually is copyright. If it is known that large parts of it were generated by LLM, they’d need evidence to demonstrate sufficient human input to establish copyrightability. If they’ve kept highly detailed traces of the development process, that could be rather straightforward; if they haven’t, it could be really difficult.

Now, that’s true in the US, which never accepted mere “sweat of the brow” as a basis for copyright; the UK courts have, and most of the Anglosphere follows the UK on this more than the US.

The other factor: when dealing with an (almost) trillion dollar corporation, even if you’ll win the legal argument, they may bankrupt you with legal fees before the argument is ever properly heard.

But I suspect the precedents on this topic are going to be established by lawsuits involving far smaller actors.

(IANAL and I speculate only for myself, not any present, past or future employers.)

> The code is not eligible for copyright.

This is very much not what the linked case established.

According to the link:

"The US Copyright Office and federal courts require human authorship for copyright protection; works created solely by AI are not eligible for registration under the current rules."

The Supreme Court declined to consider a challenge to this rule, and so for the moment at least, the rule remains in place.

This means that companies leaning heavily into their LLM use may very well find that they do not, under the law at least, actually own any of their code. As I've read elsewhere there's every possibility that AI code will be the asbestos of the Software Engineering world. Something we'll be trying to get rid of for decades, once everyone comes to their senses.

I think the word "solely" is going to be a tunnel you can drive freight trains through.

So with the asbestos analogy, we encase the fibers in resin and call the whole thing copyrighted.

Please explain. That is exactly what the linked case established.
It established you can't assign copyright to the LLM itself. That's very different.
> If they are so adamant on protecting their IP,

What they are trying to protect doesn't qualify as intellectual property. Only 4 categories of IP exist: (1) copyrights; (2) patents; (3) trade secrets; (4) trademarks.

The capabilities embedded in model outputs don't qualify. Machine-generated outputs are ineligible for copyright. They aren't covered by patents. They aren't trade secrets, because the model companies are selling them rather than keeping them secret. And of course, trademarks are conceptually inapplicable.

This leaves the model companies with contract law (ToS) which is pretty inept because it can't bind third parties. And technical measures, like the ones being discussed in the article. And, of course, politics.

Frankly, I think it's pretty ridiculous to even think that models can be protected from being learned from. I feel the Stanford Alpaca team demolished that idea 3 years ago.

The hypocrisy of the pro-AI mega corp arguments makes my head spin. For three years they’ve been using the example of a human reading books and then outputting creative works influenced by them as analogy for training AI on copyrighted works. Now suddenly we’re supposed to not draw the same parallel about a hypothetical person who learned from Claude and is now outputting creative work based on it.
> So, why all this "effort" to protect the model?

Because it's their model and business and they are free to use the free market to do exactly that?

That's their free market rights too. If you don't like it, use another model (which they would be fine with).

> Because it's their model and business and they are free to use the free market to do exactly that?

I mean, nothing stops distillers to find better ways to distill, either. Meaningless cat & mouse games.

> If you don't like it, use another model (which they would be fine with).

Thanks, I use none. It's peaceful this way.

The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not, which is what China has been doing. Countless thousands of separate requests abusing the service (which is not a simple static HTML feed, but an AI service request) for every kind of query to soak up the results.

The post is about what's in the local code, but for a long time there has already been modifications made to the request outputs from the major cloud services as they work together to both curb adversarial distillation and to degrade the quality of training China can get from that distillation.

It's likely not to make the answer wrong or bad, but to make it so that any model trained on the output would not gain the benefit of the model's reasoning generalization skills as easily and also identifying markers that might even link back to request IDs.

The techniques talked about in this post are naive and simplistic, largely because they are released publicly.

It's not as much about protecting IP as much as it is about slowing China down or being able to track the effects of abuse. So many people are talking about greedy company this, greedy company that. The world is not made up of caricatured giant money pigs wearing suits with monocles and gold watches. That is a children's view of Marx's exaggeration on free markets. Bad, greedy people do exist, but if that is your only hammer for every nail then you have a problem.

The upper-bound for how good these models can be is so crazy that it is essentially dual-use military applicable to an extent most other technologies are not. It's not only cyber attacks or biological weapons. Most people are not even built to understand the possible threats.

Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.

> Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.

Reminds me of this comic: https://xcancel.com/tomgauld/status/571994690289061888?lang=...

None of the superpowers in this world is innocent, and like MAD, more countries have the capability, the better.

I know some of the things CCP do/did. I know some of the things US does/did. I'm from neither, so I don't take sides.

AI's use has been confirmed, or more precisely boasted by two countries in two different wars, and China was not one of these countries.

We have seen the effects of "if they don't know them, they can't exploit them" mindset of NSA for years. Keeping information/technology private is neither beneficial, nor possible. It's only a temporary moat-ish gap. Not a definitive solution.

Certainly the world is full of actions and reactions, nothing is happening in a vacuum. You don't have to be from a country to take sides, but presumably you have some kind of moral compass, some kind of values around personal freedom or the worth of a human life.

There can be a very real cost, because one side comes from an ideology with a history that wants to conquer the entire Earth which caused World War 2 while the other side is trying to prune the planet like a bonsai to prevent it from descending into total chaos to preserve some sense of international order.

Europe was constantly at war, and we helped stabilize it. Middle East as been constantly at war, and if Iran can be sorted then it will be the closest to some sense of peace it's been in a long time.

We used to be in Japan, Philippines, Germany, Vietnam, South Korea, Iraq, Afghanistan and so on. How many are US territories? None. We aren't out there to conquer the globe and take land. We're usually fighting other people's wars for them, because they're up against better resourced opponents. Meanwhile China is over there building artificial islands, ramming other country's ships, creating ideological police stations in countries around the world to harass people and engaging in the most widespread international interference campaigns in human history.

They do not treat their people well and they do not have free speech. The internet is flooded with their propaganda now, because they have a human numbers advantage.

It's true that given time most advantages are temporary, but there's always that slim chance we could slow them down until the CCP collapses and they could become a more normal country.

You sound like you've swallowed pro-US-propaganda hook, line and sinker.

The reason the middle-east is at constant war is because colonialist machinations. Same goes for south-Saharan Africa. And the US is a big colonialist player, just ask Vietnam, South-America, Iran, Afghanistan, etc. They all have been attacked by the US because of US colonial interests. If anything, one could make the argument that the PRC is treading much more lightly than the US.

That said - I'm not defending the PRC by any way; it's a state-capitalist hell hole that's suppressing workers by denying them any ability to organize and whose political class is purely focused on furthering their own interests and that of the moneyed elite, the common person be damned.

The thing is - so is the US.

If you think those were about colonization, you would be well served to examine history closer.
Ah yes. The exceptionalism argument.

There's the good "us" and the bad "them."

I didn't make an exceptionalist argument, but if any country's behavior and values can be measured compared to others, you will always be able to make some kind of decision about where those fall in terms of goodness or badness. Do you not believe in good and bad?
> The CCP is darkside material.

And which country has "Black Sites" peppered around the world to detain and interrogate people they don't like?

The US doesn't have black sites anymore and when it did, the interrogation techniques were chosen to avoid physical harm. The results were bad, we didn't like it here in the US even if they were extreme measures for extreme times and so we shut it down. It had a high error rate and generally didn't reflect what we thought was right.

Meanwhile the CCP regularly abducts its own citizens and executes more people than the entire world combined.

> The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not

This is literally what the "training AI on copyrighted works is just like a human learning/getting inspired" crowd has been arguing though.

Literally. People have been literally saying that it was wrong because they did this "learning" at scale in a dishonest way.

In some ways it's an offshoot of the honest benefit of search engines already crawling all this content. That has its own conflicts, like just how much of a page's content should you reproduce in the results before it's basically considered stealing their content without benefiting the site itself.

There is a balance to strike, both in search engine fair use cases and AI fair use cases. The major cloud LLMs do double as web search engines now, though they didn't originally. In many cases there's no reason left to click the links they sourced from.

That is a legitimate concern. At least within the US, I think there are nuances around fair use and contract law. A lot of companies are getting paid for having their content used in these models, but many websites had no particular rules you had to abide by and the content was simply public. I think if you're operating under an agreement, then even if there is fair use or public domain content being reproduced by the site you are still bound by that agreement.

Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used. These AI services are obviously very different, but there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage.

I'm not exactly comfortable with the mass scale that everything was soaked up to train these models even within the umbrella of search services, but I also admit that a lot of the usage was probably quite legal. The potential displacement caused by the resulting trained models on artists or writers is almost its own facet. In practice, whether they ONLY trained on strictly legally acquired fair use content with no errors and paid agreements to acquire even more content than they already do or not, there was enough legally accessible information for fair use that there was no escaping some kind of impact on artists, writers, etc.

With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

Well it was also problematic when the search engines started quoting the websites in such a way to disincentivize people from visiting the actual website.

> At least within the US, I think there are nuances around fair use and contract law.

The concept of "fair use" as it exists in the US-law system is completely dysfunctional (see e.g. nearly every educational music channel on YouTube), so utterly biased to favour large corporations, that there's very little room for whatever "nuances" you believe exist.

> Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used.

Yes 300 year old paintings are public domain. Indeed there are certain rules for the people/institutions who digitize them. It's not "they have some say", there's actually nothing mysterious about it and it is not similar to Anthropic's copyright heist at all because nearly all of the books they copied were not more than 100 years old.

> there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage

well where I live, there are laws about what a "service" can claim to "lay out as acceptable usage" instead of the other way around ...

> I also admit that a lot of the usage was probably quite legal

Let's disagree on that. I think it wasn't a lot and the vast majority was not legal. How do you think the LLMs "learned" to speak all these non-English languages? Unless your point is that it's probably quite legal to treat foreign IP like that. Which it may very well be in the US, especially if the corporation is large enough, but imvho it's still wrong.

> With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

And with any bad luck, these AI corporations will hold frontier models hostage for the rest of time.

I honestly don't want to put that up to "luck".

oh no, the company that illegally used every possible media they could get their hands on is crying that some other company is doing something potentially shady but not illegal? And using that excuse to put in place hidden surveillance systems on their customers?
People keep throwing this idea around haphazardly, but U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use. You may not like it, but that doesn't make it "illegal".
You have to admit that "downloading every book ever written for free from a repository of books that is itself illegal to compile and to run, in order to write a text generation tool" being legal is at least unintuitive, to put it mildly.
It wasnt, that's why they paid a >billion dollar settlement over it, and now license/purchase them. I don't know if the people distilling are licensing those books/etc today, though
I'd appreciate if the down voters explain why. I wasn't making a value judgement.

Anthropic did pay more than a billion: https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...

And is now buying up a lot of books (controversially, as scanning involves cutting their spines) because that's what the law deems the legal method: https://www.washingtonpost.com/technology/2026/01/27/anthrop...

We know that models like Deepseek are trained on copyrighted books too: https://arxiv.org/abs/2603.20957

The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.

Clearly paying that fine didn't do anything to stop Anthropic from doing it again.

Buying a book doesn't make it legal to publish lossy compressed copies of it.

Also, the vast majority of authors whose work was copied against their wishes didn't receive any of that fine.

It sounds like your argument is that they paid a fine for breaking the law, and therefore it is okay they reap the benefits of breaking the law and are allowed to continue to do so?

> The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.

UHmmm you remember when Sam Altman changed his profile pic to look like a Disney version of his own face? Yeah neither do I.

Clearly US AI models are playing loose with the use of overseas IP just as much, and even publicly flaunting it, as if US-based IP is more worthy of protection but Gibli can suck it.

No it's not unintuitive.

Just like I can learn from a book and nobody can make that illegal, so can other people transformative do the same with computers.

Fair use is fair use.

Just like these distillers can learn from Claude’s output. Fair use is fair use.
I don't think Anthropic argues that distillation violates copyright. AFAIK, their position is that it violates their terms and conditions for interacting with their servers.
For someone who finds this "not unintuitive" you sure are confused!

"Just like I can learn from a book" - ok. Are you allowed to go to libgen and download a book in order to learn from it, because learning is a fair use?

Has it? Because as far as I can tell those cases keep getting settled out of court before a legal precedent can be set.

For record breaking amounts too.

Maybe "U.S. courts have pretty consistently decided" used to mean something, but I don't think the opinion of US courts should be the standard for anything, anymore.
> U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use.

I don't believe that this has been resolved at all, and there are quite a few pending lawsuits about it at this very moment.

The courts have never said piracy, which is how the training sets were originally built, is legal. There are several court cases still ongoing over this.
Right, so it seems that distilling an AI model is legal too then. At least it is somewhat similar.
Legal vs "They aren't going to let you do it with their service" are two different things.
Screw those poor copyright holders without the means to stop frontier AI labs, amirite?
>Screw those poor copyright holders

In general yes. Cut it down to a reasonable amount of time and I'll care a whole lot more about those 'rights' holders.

It is a violation of their terms of service.

There are plenty of good reasons to not use Anthropic's services. If you don't like their terms of service, do stop using them! I personally think Anthropic's increasingly successful attempts at regulatory capture are even more distasteful.

Oh Anthropic has shown their ugliness in more ways than one I agree. You have to have to done some pretty heinous shit for openAI to look good in comparison.
It was also a violation of the terms of service of those books (aka copyright)
> that training on [lawfully obtained] copyrighted works falls under fair use

Fixed that for you.

Were the copyright owners contacted prior to this lawful obtaining that you speak of? Or after?
I miss the days when tech people were copyright skeptics. Remember when everyone was upset with Disney for our perpetual copyright regime and the destruction of public domain?

Now many tech people are copyright maximalists and 100% converted to the church of Disney. It’s depressing.

I don't think that's right. The problem is that Anthropic is hoarding it and that's hypocritical. If copyright doesn't count for Anthropic, they should publish Claude. If they wanna hide Claude behind copyrights and/or TOS, they don't get to screw with other people's copyrights and TOS and then profit from it.

To call that opinion "copyright maximalist 100% converted to the church of Disney" is, at the very least, hyperbole.

its not copyright maximalism. people just see the obvious hypocrisy. a lot of people are also fine with some copyright
> loudly stating the foreign labs have been distilling their models for a while now.

They would be stating this even if it weren't true, because it fits their marketing.

While I don't disbelieve the claim outright, I highly suspect Anthropic is misleading everyone about the severity.

Distillation usage still burnishes usage numbers for IPO...

If anything, Anthropic is incentivized to track but do nothing until equity lock up expires.

No sympathy for them trying to “protect” the output of a model that’s trained on data that they didn’t get consent to use. Ripe hypocrisy.
> foreign labs

Apparently not just foreign labs. It looks like xAI distilled Anthropic models to train grok.

https://opentools.ai/news/xai-trained-coding-models-claude-o...

That's less of a worry though since xAI is patently incompetent.
Incompetence is not an excuse for amorality
Oh that’s ok, xAI is enthusiastically immoral.
What's amoral about distillation?
Except when it comes to image models. Imagine is extremely good and extremely cheap. I've been using it to generate book covers for ebooks (old novel short stories that never had a cover, for example) and it's phenomenal. Each cover is about 6 cents
I wouldn't be surprised
I really doubt other labs are distilling Claude using the Claude Code CLI when they can way more easily use the API directly.

I also don’t get why the « protection » on ANTHROPIC_BASE_URL. If I change it to use a Chinese model, the Chinese model will not care at all about the modified prompt. On the contrary, if I’m distillating (which again, using CC CLI would be stupid), I’m not going to change ANTHROPIC_BASE_URL.

Sounds suspiciously similar to the "album title", "Steal this album" by system of a down.

Im not sure why we are dithering on the boundaries of honesty when the entire content LLMs are trained on is stolen.

Are we debating "honor among thieves"?

Of course we are not, or maybe we are!

Does the behavior of a thief even matter to me? only after they do their time. And they will.

I can see the investors perched on the balconies of their condos in a couple years if that.

its a long way down.

"Steal this book" by Abbie Hoffman
The obvious response is the realization that spending trillions on training LLMs is not a viable business model if they can be distilled for a much lower cost.
What does that have to do with CC? I'm not commenting on that being good/bad/legal/illegal, but CC is separate from the models. If they really are doing this maliciously it is because they are trying to ignore my 'CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1' flag (if that still means anything).
The article seems to state as much minus the obfuscation. However justified they are to respond, this can be a slippery slope. We're bound to hear more reports of hidden user data exfiltration.
The have been fucking distilling our websites and writing, even when behind TOS, aggressively bypassing protection mechanisms. They they can fuck right off
Aww man that's rough did someone steal their content and use it to make money without asking them?
> a mechanism to make that obvious.

Say they prove that foreign labs are distilling their models, then what?

Something about throwing stones in glass houses.
>very loudly stating the foreign labs have been distilling their models

Help! Someone else is blatantly ripping off my plagiarism machine!

i think you're being played by their whole ethical high horse PR angle