Hacker News new | ask | show | jobs
by jackcviers3 1 day ago
On the other hand, the open weights models could crawl and annotate and rl the training data that Anthropic and OAI did in exactly the same way, and take the exact same legal hits. They use distillation because it's cheaper not to do so.
2 comments

> On the other hand,

What is the other hand here? They could do slow/costly illegal think #1 instead of the fast/cheaper illegal thing #2 that they currently do?

Neither one is illegal though. Scraping the internet is as legal as surfing the internet. And what Anthropic calls "distillation attacks" is really just paying for and using the service Anthropic provides. I would think a judge would have the same view of distillation as they do of scraping, if Anthropic doesn't want a subscriber to have access to the model, they are within their rights to block access, but it's not the user's job to refrain from using the service. If their usage is so different from everyone else's, they should be easy to detect and block. If their usage is so similar to everyone else's that it's difficult to detect, then it shouldn't be of any concern.
Genuine question: is distillation illegal, or just against Anthropic's terms of use?
It seems to be only against ToS; the problem, however, is that “distillation” involves intent. Anthropic seems to be calling for ban on distillation because they cannot reliably identify who does it - if they could, they would’ve just banned all infringing users.

However, if there is no systematic difference between regular use and “distillation use” then they need to establish intent, which is laughably hard if you cannot even distinguish use cases.

Hence the call to make it illegal too (and just against ToS). Because if it is illegal, then they can rely on US to enforce the law. And the US, for a lack of technical solution, can use the good old “China bad” shortcut.

Hmm, why do you think this is true? One reason I'm skeptical of this is because RL envs are often purchased (and are not publicly available), and this might be a sizable component of why models are getting better.