Hacker News new | ask | show | jobs
by theplumber 16 days ago
>> No they're not. It would end both companies if they were ever found to be doing that. Their terms are clear - The argument here is that with the Chinese labs you have zero legal recourse.

Their terms are not worth shit considering they are reselling you stolen copyrighted data. Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want. Perhaps you didn't watch the space with Anthropic going back and forth with ToS updates(we retain your data for 30 days...stike that and add 30 days or more or no or ..whatever) like my own alpha website.

6 comments

There is an enormous difference between:

* Exploiting ambiguity around fair use at a large scale before the law catches up and then jointly lobbying with your competition to make sure your interpretation of the law becomes reality.

* Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

The former is firmly in the gray area of legality and doesn't directly hurt your own customers. The latter is both an unambiguous contract violation and a flagrant attack on your own customers' most valuable asset.

https://www.anthropic.com/legal/privacy

> Personal data we collect or receive to train our models

> • Data that our users or crowd workers provide, including Inputs and Outputs from our Services (unless users opt out)

> • Feedback that users explicitly provide about our Services

> • Materials flagged for safety, security, or policy review

While I don’t have visibility into individual corp contracts, hitting tab on a FIM is ‘feedback’, so it is not so clear cut.

First: This is the general privacy policy, not the enterprise contract. I don't know what goes into the enterprise contract, but I do know that our legal department spent a very long time making sure it was satisfactory before we got access.

Second: My argument doesn't hinge on Anthropic not being able to weasel their way out in court if it came to that. My argument is that neither Anthropic nor OpenAI are going to break their signed contracts or even fudge on the clearly communicated understandings of what the terms of the API pricing are because neither one wants to hand the other the obvious weapon of: "unlike {other guys} we honor our word".

It's just not happening, and comparisons upthread to the fair use story totally misunderstand the incentives at play here.

(And as an aside, this whole thread also shows clearly the classic programmer misunderstanding of the law. The peanut butter sandwich instructions analogy is for code, not for the law. The law doesn't actually work by allowing any possible interpretation to hold equal weight the way that many programmers think it does.)

Moonshot says the same thing, that if you don't want to be trained on, get an enterprise contract.
> The law doesn't actually work by allowing any possible interpretation to hold equal weight the way that many programmers think it does

Is that so? Recent rulings in the US specifically gave me the impression that when backed by sufficient legal representation and goodwill on the judging side indeed any possible interpretation will suffice.

I think that's what makes law making complicated - you either err on the side of leaving too much room for interpretation or not enough.

> Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

You mean all the conditions that are attached to Fable use? My enterprise is deliberately holding off because those are unacceptable.

That suggests the system working as it should. They present terms for use, you don't like them, so your don't use it. One of their other products has terms you're ok with, so you use that product.

Good, fine. This is an example of trusting the company to honor their own terms, not the opposite.

That feels like moving the goalpost. First they would "never not respect the enterprises IP" then the next message "Oh, but it's fine as long as they introduce new terms that you can reject"

Either they respect IP, or they don't. Clearly they don't.

retention for 'safety' -> AI race as national security -> training on your data for 'national security' aka safety

It's simple mental calisthenics. If you are handing an organization whose entire business model is built on stealing data with spurious reasoning, what do you actually expect they will do? Don't be a fool.

This argument would hold more weight if Anthropic and OpenAI main customers weren’t massive trillion dollar companies with legal teams capable of burying just about anyone, anywhere, for even the mildest contract violation. Something that OpenAI is getting some close up experience with at the moment.
I'd like to see you try using mental calisthenics against a well-funded legal department. Let me know what the judge says.
Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

I'm not comfortable with how these models were trained. I have quite a bit of open source code out there, and I personally see such training as copyright and license laundering.

But that's not how the law sees it, and I grudgingly accept that, regardless of how I may feel, and I don't let my feelings on the matter make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Sure, they might be breaking their promises, training on our data when they say they won't. But I do think they most likely aren't, and that it would be corporate suicide if they were and it ever came out.

> But that's not how the law sees it

Anthropic paid several billion dollars to settle a lawsuit they were likely to lose. OpenAI is now about to get taken to the cleaners for corporate espionage against Apple. They do not give a fuck about the law. Paying $5 billion for some fines is a trivial cost of doing business when you're aiming for trillion-dollar IPOs.

> make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Irrationality is thinking there's such a thing as honor and that companies which have repeatedly broken the law for data won't do it again when there's no enforcement mechanism that acts as a real deterrent.

Honour is not the relevant point. What's relevant is that breaking your TOS would give many large individual companies (a) a clear case to sue you and make a lot of money (b) a reason to avoid using you ever again, because who wants their data leaking into a potential competitor? Scraping the internet, as well as being much harder to adjudicate legally, also creates far fewer powerful companies with means, motive and opportunity to take you to court.
Their argument boils down to "they've done it once and nobody prevents them from doing it again"

This dog-and-pony-show is a rehash of the Pascal's wager we saw with smartphone security. Everyone thought it would be "corporate suicide" to hack an iPhone, but NSO Group did it. Apple sued NSO Group, and then settled out of court immediately after. Now we live in a post-hacking world and everyone pretends like this is an unavoidable necessary evil that corporations are powerless to stop. Suggesting litigation is a comically useless strategy because the law rubberstamps any form of useful surveillance or retention. Failing that, NSO Group has enough sycophant lobbyists to smear anyone that takes their threat seriously. Look at OpenAI and Anthropic and tell me that it's not the same hostage situation; can you?

You can do whatever stupid stuff you want to with your data. But this is an absurd amount of faith to give to guilty businesses, on the level of planning your world domination schemes over Skype.

>> Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

Not at all. My point is that the every thing they do is quite questionable from business development to sales & marketing

> But that's not how the law sees it, and I grudgingly accept that

I think that is sort of their point. There was one thing that you, I, and millions of others would call infringement, (scraping the whole Internet to train proprietary models) but the law deemed it "fair use", and they got away with it with impunity. Now there is this other thing that we'd all (easily) call infringement, and I understand why people doubt that this time will be any different.

Anthropic paid a large settlement for the copyrighted data they pirated. So far, US courts have found that it's perfectly fine to train AIs on copyrighted data for which you have legal access.
I’ve always considered this a token gesture. They paid $3000 each to 500,000 authors. Doesn’t change the fact that each author’s blood sweat and tears were the input to their machine, and they can make money on the output of that machine in perpetuity.
But if Anthropic had just bought the authors' books instead of pirating them, then they would have paid nothing.
> Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want.

The discussion was about training, not data retention. Two very different concerns.

And if you're a decent sized customer, most providers have a route to not even retaining the data for safety/security reasons. The reason Anthropic had issues is because they do have a path to "no data storage" for Sonnet/Opus, but not for Fable. Which is why at work we have access to the former, but not the latter.

while it's plausible that Sam Altman could find someone to covertly exfiltrate privileged data and then somehow covertly train on it while the rest of the developers remain ignorant it would all come to naught when the companies from whom they stole data probe the model with questions it should not be able to answer.

I don't believe anyone knows how to train the model in such a way that it's guaranteed not to remember any specifics while still having the training run be worth anything.

That's their problem.
Whether the terms are worth shit doesn't matter. If they're training on data from paying customers who have requested otherwise and it gets out (which it would, eventually), SAP, Accenture, Deloitte and other huge companies with well-funded legal teams would nuke them from orbit. This is a different area of law from the copyright stuff, different rules/norms/expectations/consequences apply.
They're not training on your data, they're training on "please anonymise this conversation" data.
So because it would wreck they if others found out, it’s unlikely?

Which is more likely? That past behavior is an indication of future behavior, or that they because they could be eliminated from being found out it’s unlikely they’d do that thing. (By the way it’s also likely they’d are eliminated if they dont train their data with every advantage over their competitors possible). So I think it’s naive to think the incentives reward not doing the malicious thing now.