Hacker News new | ask | show | jobs
by nikcub 12 days ago
> I pretty sure OpenAI and Anthropic are doing the same or worse.

No they're not. It would end both companies if they were ever found to be doing that.

Their terms are clear - if you use the coding plans they can[0] train in return. Enterprise and API, absolutely not.

The argument here is that with the Chinese labs you have zero legal recourse.

[0] opt-in, thanks

17 comments

>> No they're not. It would end both companies if they were ever found to be doing that. Their terms are clear - The argument here is that with the Chinese labs you have zero legal recourse.

Their terms are not worth shit considering they are reselling you stolen copyrighted data. Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want. Perhaps you didn't watch the space with Anthropic going back and forth with ToS updates(we retain your data for 30 days...stike that and add 30 days or more or no or ..whatever) like my own alpha website.

There is an enormous difference between:

* Exploiting ambiguity around fair use at a large scale before the law catches up and then jointly lobbying with your competition to make sure your interpretation of the law becomes reality.

* Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

The former is firmly in the gray area of legality and doesn't directly hurt your own customers. The latter is both an unambiguous contract violation and a flagrant attack on your own customers' most valuable asset.

https://www.anthropic.com/legal/privacy

> Personal data we collect or receive to train our models

> • Data that our users or crowd workers provide, including Inputs and Outputs from our Services (unless users opt out)

> • Feedback that users explicitly provide about our Services

> • Materials flagged for safety, security, or policy review

While I don’t have visibility into individual corp contracts, hitting tab on a FIM is ‘feedback’, so it is not so clear cut.

First: This is the general privacy policy, not the enterprise contract. I don't know what goes into the enterprise contract, but I do know that our legal department spent a very long time making sure it was satisfactory before we got access.

Second: My argument doesn't hinge on Anthropic not being able to weasel their way out in court if it came to that. My argument is that neither Anthropic nor OpenAI are going to break their signed contracts or even fudge on the clearly communicated understandings of what the terms of the API pricing are because neither one wants to hand the other the obvious weapon of: "unlike {other guys} we honor our word".

It's just not happening, and comparisons upthread to the fair use story totally misunderstand the incentives at play here.

(And as an aside, this whole thread also shows clearly the classic programmer misunderstanding of the law. The peanut butter sandwich instructions analogy is for code, not for the law. The law doesn't actually work by allowing any possible interpretation to hold equal weight the way that many programmers think it does.)

Moonshot says the same thing, that if you don't want to be trained on, get an enterprise contract.
> The law doesn't actually work by allowing any possible interpretation to hold equal weight the way that many programmers think it does

Is that so? Recent rulings in the US specifically gave me the impression that when backed by sufficient legal representation and goodwill on the judging side indeed any possible interpretation will suffice.

I think that's what makes law making complicated - you either err on the side of leaving too much room for interpretation or not enough.

> Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

You mean all the conditions that are attached to Fable use? My enterprise is deliberately holding off because those are unacceptable.

That suggests the system working as it should. They present terms for use, you don't like them, so your don't use it. One of their other products has terms you're ok with, so you use that product.

Good, fine. This is an example of trusting the company to honor their own terms, not the opposite.

That feels like moving the goalpost. First they would "never not respect the enterprises IP" then the next message "Oh, but it's fine as long as they introduce new terms that you can reject"

Either they respect IP, or they don't. Clearly they don't.

retention for 'safety' -> AI race as national security -> training on your data for 'national security' aka safety

It's simple mental calisthenics. If you are handing an organization whose entire business model is built on stealing data with spurious reasoning, what do you actually expect they will do? Don't be a fool.

This argument would hold more weight if Anthropic and OpenAI main customers weren’t massive trillion dollar companies with legal teams capable of burying just about anyone, anywhere, for even the mildest contract violation. Something that OpenAI is getting some close up experience with at the moment.
I'd like to see you try using mental calisthenics against a well-funded legal department. Let me know what the judge says.
Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

I'm not comfortable with how these models were trained. I have quite a bit of open source code out there, and I personally see such training as copyright and license laundering.

But that's not how the law sees it, and I grudgingly accept that, regardless of how I may feel, and I don't let my feelings on the matter make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Sure, they might be breaking their promises, training on our data when they say they won't. But I do think they most likely aren't, and that it would be corporate suicide if they were and it ever came out.

> But that's not how the law sees it

Anthropic paid several billion dollars to settle a lawsuit they were likely to lose. OpenAI is now about to get taken to the cleaners for corporate espionage against Apple. They do not give a fuck about the law. Paying $5 billion for some fines is a trivial cost of doing business when you're aiming for trillion-dollar IPOs.

> make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Irrationality is thinking there's such a thing as honor and that companies which have repeatedly broken the law for data won't do it again when there's no enforcement mechanism that acts as a real deterrent.

Honour is not the relevant point. What's relevant is that breaking your TOS would give many large individual companies (a) a clear case to sue you and make a lot of money (b) a reason to avoid using you ever again, because who wants their data leaking into a potential competitor? Scraping the internet, as well as being much harder to adjudicate legally, also creates far fewer powerful companies with means, motive and opportunity to take you to court.
Their argument boils down to "they've done it once and nobody prevents them from doing it again"

This dog-and-pony-show is a rehash of the Pascal's wager we saw with smartphone security. Everyone thought it would be "corporate suicide" to hack an iPhone, but NSO Group did it. Apple sued NSO Group, and then settled out of court immediately after. Now we live in a post-hacking world and everyone pretends like this is an unavoidable necessary evil that corporations are powerless to stop. Suggesting litigation is a comically useless strategy because the law rubberstamps any form of useful surveillance or retention. Failing that, NSO Group has enough sycophant lobbyists to smear anyone that takes their threat seriously. Look at OpenAI and Anthropic and tell me that it's not the same hostage situation; can you?

You can do whatever stupid stuff you want to with your data. But this is an absurd amount of faith to give to guilty businesses, on the level of planning your world domination schemes over Skype.

>> Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

Not at all. My point is that the every thing they do is quite questionable from business development to sales & marketing

> But that's not how the law sees it, and I grudgingly accept that

I think that is sort of their point. There was one thing that you, I, and millions of others would call infringement, (scraping the whole Internet to train proprietary models) but the law deemed it "fair use", and they got away with it with impunity. Now there is this other thing that we'd all (easily) call infringement, and I understand why people doubt that this time will be any different.

Anthropic paid a large settlement for the copyrighted data they pirated. So far, US courts have found that it's perfectly fine to train AIs on copyrighted data for which you have legal access.
I’ve always considered this a token gesture. They paid $3000 each to 500,000 authors. Doesn’t change the fact that each author’s blood sweat and tears were the input to their machine, and they can make money on the output of that machine in perpetuity.
But if Anthropic had just bought the authors' books instead of pirating them, then they would have paid nothing.
> Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want.

The discussion was about training, not data retention. Two very different concerns.

And if you're a decent sized customer, most providers have a route to not even retaining the data for safety/security reasons. The reason Anthropic had issues is because they do have a path to "no data storage" for Sonnet/Opus, but not for Fable. Which is why at work we have access to the former, but not the latter.

while it's plausible that Sam Altman could find someone to covertly exfiltrate privileged data and then somehow covertly train on it while the rest of the developers remain ignorant it would all come to naught when the companies from whom they stole data probe the model with questions it should not be able to answer.

I don't believe anyone knows how to train the model in such a way that it's guaranteed not to remember any specifics while still having the training run be worth anything.

That's their problem.
Whether the terms are worth shit doesn't matter. If they're training on data from paying customers who have requested otherwise and it gets out (which it would, eventually), SAP, Accenture, Deloitte and other huge companies with well-funded legal teams would nuke them from orbit. This is a different area of law from the copyright stuff, different rules/norms/expectations/consequences apply.
They're not training on your data, they're training on "please anonymise this conversation" data.
So because it would wreck they if others found out, it’s unlikely?

Which is more likely? That past behavior is an indication of future behavior, or that they because they could be eliminated from being found out it’s unlikely they’d do that thing. (By the way it’s also likely they’d are eliminated if they dont train their data with every advantage over their competitors possible). So I think it’s naive to think the incentives reward not doing the malicious thing now.

I would think they are not but Alex Karp CEO of Palantir seems to imply that they are:

https://youtu.be/0A3sGymV6kY?si=ti7uSZtYqJ3vKpGM

I found it a little shocking TBH

Alex Karp says a lot of things
> if you use the coding plans they train in return.

No, you have to opt-in to that. There's a privacy toggle on account settings.

For OpenAI, you have to use the Enterprise plan at API pricing in order for them not to train on your data.

Source: https://chatgpt.com/codex/pricing/?type=team

Not true. You can turn off training in settings on an individual plan.
Yes, my comment only applies to Claude - I have no knowledge of OpenAI policy.
Are we talking about the company sending back private information through its client to « fight » model distillation?
Yes.

Enterprise contracts are checked and agreed by lawyers. The contract states no training.

If the provider fucks up, there are actual monetary damages defined for breach of contract.

It's an unenforceable clause. The affected party has no means to prove that a breach has happened.
I think we all ought to look at the ZDR fine-print here.

I get that in principle that there's no retention, but these are powerful models that can comprehend, paraphrase and summarize your logs for the sake of "product" improvement. Who knows what's collected here.

> It's an unenforceable clause. The affected party has no means to prove that a breach has happened.

Big tech spends hundreds of millions in high powered lawyers, audit logs, and contractual agreements with the sole purpose of proving your point wrong.

Cool. And have they?
Anthropic (and probably OpenAI) can technically flag anything for 'safety' and view it, I'm not trying to imply that they are doing it arbitrarily, but they do have a way to do it within the ToS as far as I understand (happy to be proven wrong).
Are there terms clear? I dunno. There are ways to train on API usage without training on API usage.
How would you explain how they built https://github.com/anthropics/claude-for-legal then?

> ... for the legal workflows we see most

see.. where?

claude.com / chatgpt.com. By default your chats over there are opt-in for "help improve our models".

Not the API.

Terms for the app: https://privacy.claude.com/en/articles/10023580-is-my-data-u...

The API: https://privacy.claude.com/en/articles/7996868-is-my-data-us...

If you are a regular user they could care less,if you are enterprise they might be more careful. They have credit card info, and all your chats so I’m sure they can figure things out
Maybe Apple lawyers can figure this out in discovery.
If that was completely the whole story, then the extra Zero Retention etc tier enterprise things wouldn’t exist

I also don’t trust them lol

How much do we trust those guarantees? Apple currently alleges that OpenAI employees have been stealing Apple's trade secrets: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secre...

> [Mr. Tan] has directed job candidates still working for Apple to bring “Actual parts” from Apple to their interviews for “show and tell” sessions in which he and his team at OpenAI can elicit still more Apple confidential information.

> As part of its investigation, Apple found a “pattern by employees who depart for OpenAI of taking steps to evade the security processes intended to protect Apple’s confidential information.”

> Apple also claims former engineer Liu exploited a security bug to download confidential engineering files after leaving the company. Rather than report the exploit, Liu allegedly joked about it in messages (“LOL,” “so funny”). Liu also failed to return an Apple-issued laptop after his departure.

This seems pretty close to "they trust me, dumb fucks" behaviour.

They could use an agent to summarise the source material, and then train models on those summaries, and claim that some sort of clean-room training has happened?
You've been to too many meetings with PMs and directors saying "An agent could very easily do this"
I understand you're trying to be funny, but my point is that with novel technology there are novel ways to claim innocence in courts because of the legislative void.
my experience right now with fable via bedrock is they collect data, 5.6 has the option for ZDR at least
I think the risk of not doing is more existential than doing so and getting caught. Wouldn’t you agree?

Edit: And the point of the poster is they have already demonstrated a track record of lying and misconduct, so how can you trust their word now? What have they done to show you they have taken responsibility for past actions and changed?

lmao, wasn't xAI caught doing this recently? moreover at least moonshot is being honest about it.
they train on your requests by paraphrasing them (which means rewriting them but keeping all the saliency) and removing their association with you

i don't know why this is so controversial, their terms are written to perfectly fit this training regime. one of you downvoters i'm sure has an enterprise contract with them, just ask.

if you are using bedrock, until very recently, they didn't see your requests and could not paraphrase. but too many people were using bedrock for too much stuff they wanted to see. so that's why the terms for bedrock changed for fable 5. this was the core of the palantir / defense dept drama with anthropic.

Anthropic constantly uses dark patterns to steal training data from customers (like the “how is claude doing” spam, data retention loosening when the safeguards false positive, etc).
How is that a dark pattern? What is the light pattern for getting feedback from users?
There are multiple ways to use feedback. Personally I would often be fine with actual human reading my feedback, taking in account the points I made and evaluating how it should affect their feature development and future roadmap.

Now what I would expect AI companies to do is to take things which were submitted as feedback and pretty much adding to training:

"Do more of this: <copy of the whole response which was flagged as good in feedback>"

"Do less of this: <copy of the whole response which was flagged as bad in feedback>"

It's paraphrased, but the point is that they will most likely use it more-or-less as-is and thus whatever is in there will be part of the model's training set rather than someone picking up the parts from response that are important and only including them (which happens with traditional feedback).

If you have a naive loop like this you will quickly poison your dataset with e.g. thumbs downs because someone is unhappy with the latest frontend update, or teach models to not correctly refuse. I’m sure the pipeline is more sophisticated and in the middle.
Sure, but the point is more that once you submit feedback then the usual "opt out of using my data for training" no longer apply and at least that reply (and possibly whole conversation) can be included in training set in one way or another.
But we already know they're not using the transcript for training (unless you opted in t that)
They don't use your general chats it if you have opted out (note: opted out, not opted in). However if you submit feedback then whole conversation can be used.

https://help.openai.com/en/articles/5722486-how-your-data-is...

> Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.

https://privacy.claude.com/en/articles/7996885-how-do-you-us...

> If you explicitly report materials to us (e.g.via our thumbs up/down feedback mechanisms), or by otherwise explicitly opting in to training, then we may use those materials to train our models.

Again, that's a yes for OpenAI and a no for Anthropic.
Not spam? There’s no “do not ask me again”.

Not typosquat? I responded with a sentence beginning in “1” once, and it jumped in during the race. It should have prompted with something like “WARNING: This will allow us to use this session including your source code for training, which is in violation of your account settings. Proceed with “Yes I understand”.