Hacker News new | ask | show | jobs
by dovin 13 days ago
Just in case you were thinking of signing up directly with Moonshot to use the service, they appear to train even on API use:

> We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI models may contact Moonshot AI to discuss available enterprise arrangements or separate written agreements. Unless otherwise expressly agreed in writing, Customer Content may be used for the foregoing purposes.

https://platform.kimi.ai/docs/agreement/modeluse#4-content

6 comments

If they want to take my awful prompts and poison their model with it, it's their loss.

For anyone else that believes their input needs secrecy, you need to check the corporate plan of any provider for data protection clauses. Most use the cheaper plans as bait to get more training data and free feedback.

Interesting. OpenRouter classifies the Moonshot provider as ZDR. I wonder whether they have a ZDR agreement or it's a misclassification on their part.
OpenRouter's ToS also seems to allow them to store your submitted prompts anyway, so privacy advocates would have to look elsewhere anyway, that's at least how I understand it (and it surprised me).
TrustedRouter (my site) has open source proof of confidential compute that we aren’t looking at prompts or output

https://trustedrouter.com/

Why risk it either way if they provide weights for others to run this?

Am I being overly cautious not wanting to send my data to Chinese companies?

Your safety is more at risk with your data in the US government's hands.
My gut feeling is that Moonshot are probably ZDR but their terms are excessively permissive.

That said, I wouldn't rule out OpenRouter misclassifying - I've seen some providers where I'm fairly sure they have.

[flagged]
> I pretty sure OpenAI and Anthropic are doing the same or worse.

No they're not. It would end both companies if they were ever found to be doing that.

Their terms are clear - if you use the coding plans they can[0] train in return. Enterprise and API, absolutely not.

The argument here is that with the Chinese labs you have zero legal recourse.

[0] opt-in, thanks

>> No they're not. It would end both companies if they were ever found to be doing that. Their terms are clear - The argument here is that with the Chinese labs you have zero legal recourse.

Their terms are not worth shit considering they are reselling you stolen copyrighted data. Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want. Perhaps you didn't watch the space with Anthropic going back and forth with ToS updates(we retain your data for 30 days...stike that and add 30 days or more or no or ..whatever) like my own alpha website.

There is an enormous difference between:

* Exploiting ambiguity around fair use at a large scale before the law catches up and then jointly lobbying with your competition to make sure your interpretation of the law becomes reality.

* Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

The former is firmly in the gray area of legality and doesn't directly hurt your own customers. The latter is both an unambiguous contract violation and a flagrant attack on your own customers' most valuable asset.

https://www.anthropic.com/legal/privacy

> Personal data we collect or receive to train our models

> • Data that our users or crowd workers provide, including Inputs and Outputs from our Services (unless users opt out)

> • Feedback that users explicitly provide about our Services

> • Materials flagged for safety, security, or policy review

While I don’t have visibility into individual corp contracts, hitting tab on a FIM is ‘feedback’, so it is not so clear cut.

First: This is the general privacy policy, not the enterprise contract. I don't know what goes into the enterprise contract, but I do know that our legal department spent a very long time making sure it was satisfactory before we got access.

Second: My argument doesn't hinge on Anthropic not being able to weasel their way out in court if it came to that. My argument is that neither Anthropic nor OpenAI are going to break their signed contracts or even fudge on the clearly communicated understandings of what the terms of the API pricing are because neither one wants to hand the other the obvious weapon of: "unlike {other guys} we honor our word".

It's just not happening, and comparisons upthread to the fair use story totally misunderstand the incentives at play here.

(And as an aside, this whole thread also shows clearly the classic programmer misunderstanding of the law. The peanut butter sandwich instructions analogy is for code, not for the law. The law doesn't actually work by allowing any possible interpretation to hold equal weight the way that many programmers think it does.)

> Explicitly signing a contract with enterprises to respect their IP and then proceeding to break that contract with your own customers.

You mean all the conditions that are attached to Fable use? My enterprise is deliberately holding off because those are unacceptable.

That suggests the system working as it should. They present terms for use, you don't like them, so your don't use it. One of their other products has terms you're ok with, so you use that product.

Good, fine. This is an example of trusting the company to honor their own terms, not the opposite.

retention for 'safety' -> AI race as national security -> training on your data for 'national security' aka safety

It's simple mental calisthenics. If you are handing an organization whose entire business model is built on stealing data with spurious reasoning, what do you actually expect they will do? Don't be a fool.

This argument would hold more weight if Anthropic and OpenAI main customers weren’t massive trillion dollar companies with legal teams capable of burying just about anyone, anywhere, for even the mildest contract violation. Something that OpenAI is getting some close up experience with at the moment.
I'd like to see you try using mental calisthenics against a well-funded legal department. Let me know what the judge says.
Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

I'm not comfortable with how these models were trained. I have quite a bit of open source code out there, and I personally see such training as copyright and license laundering.

But that's not how the law sees it, and I grudgingly accept that, regardless of how I may feel, and I don't let my feelings on the matter make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Sure, they might be breaking their promises, training on our data when they say they won't. But I do think they most likely aren't, and that it would be corporate suicide if they were and it ever came out.

> But that's not how the law sees it

Anthropic paid several billion dollars to settle a lawsuit they were likely to lose. OpenAI is now about to get taken to the cleaners for corporate espionage against Apple. They do not give a fuck about the law. Paying $5 billion for some fines is a trivial cost of doing business when you're aiming for trillion-dollar IPOs.

> make me think irrationally when it comes to whether or not these AI companies honor the terms they provide.

Irrationality is thinking there's such a thing as honor and that companies which have repeatedly broken the law for data won't do it again when there's no enforcement mechanism that acts as a real deterrent.

Honour is not the relevant point. What's relevant is that breaking your TOS would give many large individual companies (a) a clear case to sue you and make a lot of money (b) a reason to avoid using you ever again, because who wants their data leaking into a potential competitor? Scraping the internet, as well as being much harder to adjudicate legally, also creates far fewer powerful companies with means, motive and opportunity to take you to court.
Their argument boils down to "they've done it once and nobody prevents them from doing it again"

This dog-and-pony-show is a rehash of the Pascal's wager we saw with smartphone security. Everyone thought it would be "corporate suicide" to hack an iPhone, but NSO Group did it. Apple sued NSO Group, and then settled out of court immediately after. Now we live in a post-hacking world and everyone pretends like this is an unavoidable necessary evil that corporations are powerless to stop. Suggesting litigation is a comically useless strategy because the law rubberstamps any form of useful surveillance or retention. Failing that, NSO Group has enough sycophant lobbyists to smear anyone that takes their threat seriously. Look at OpenAI and Anthropic and tell me that it's not the same hostage situation; can you?

You can do whatever stupid stuff you want to with your data. But this is an absurd amount of faith to give to guilty businesses, on the level of planning your world domination schemes over Skype.

>> Your argument boils down to "they've done something I find objectionable, so that means everything they say must be lies".

Not at all. My point is that the every thing they do is quite questionable from business development to sales & marketing

> But that's not how the law sees it, and I grudgingly accept that

I think that is sort of their point. There was one thing that you, I, and millions of others would call infringement, (scraping the whole Internet to train proprietary models) but the law deemed it "fair use", and they got away with it with impunity. Now there is this other thing that we'd all (easily) call infringement, and I understand why people doubt that this time will be any different.

Anthropic paid a large settlement for the copyrighted data they pirated. So far, US courts have found that it's perfectly fine to train AIs on copyrighted data for which you have legal access.
I’ve always considered this a token gesture. They paid $3000 each to 500,000 authors. Doesn’t change the fact that each author’s blood sweat and tears were the input to their machine, and they can make money on the output of that machine in perpetuity.
But if Anthropic had just bought the authors' books instead of pirating them, then they would have paid nothing.
> Even in they terms they started clearly say they retain your data for "safety reasons" for however long they want.

The discussion was about training, not data retention. Two very different concerns.

And if you're a decent sized customer, most providers have a route to not even retaining the data for safety/security reasons. The reason Anthropic had issues is because they do have a path to "no data storage" for Sonnet/Opus, but not for Fable. Which is why at work we have access to the former, but not the latter.

while it's plausible that Sam Altman could find someone to covertly exfiltrate privileged data and then somehow covertly train on it while the rest of the developers remain ignorant it would all come to naught when the companies from whom they stole data probe the model with questions it should not be able to answer.

I don't believe anyone knows how to train the model in such a way that it's guaranteed not to remember any specifics while still having the training run be worth anything.

That's their problem.
Whether the terms are worth shit doesn't matter. If they're training on data from paying customers who have requested otherwise and it gets out (which it would, eventually), SAP, Accenture, Deloitte and other huge companies with well-funded legal teams would nuke them from orbit. This is a different area of law from the copyright stuff, different rules/norms/expectations/consequences apply.
They're not training on your data, they're training on "please anonymise this conversation" data.
So because it would wreck they if others found out, it’s unlikely?

Which is more likely? That past behavior is an indication of future behavior, or that they because they could be eliminated from being found out it’s unlikely they’d do that thing. (By the way it’s also likely they’d are eliminated if they dont train their data with every advantage over their competitors possible). So I think it’s naive to think the incentives reward not doing the malicious thing now.

I would think they are not but Alex Karp CEO of Palantir seems to imply that they are:

https://youtu.be/0A3sGymV6kY?si=ti7uSZtYqJ3vKpGM

I found it a little shocking TBH

Alex Karp says a lot of things
> if you use the coding plans they train in return.

No, you have to opt-in to that. There's a privacy toggle on account settings.

For OpenAI, you have to use the Enterprise plan at API pricing in order for them not to train on your data.

Source: https://chatgpt.com/codex/pricing/?type=team

Not true. You can turn off training in settings on an individual plan.
Yes, my comment only applies to Claude - I have no knowledge of OpenAI policy.
Are we talking about the company sending back private information through its client to « fight » model distillation?
Yes.

Enterprise contracts are checked and agreed by lawyers. The contract states no training.

If the provider fucks up, there are actual monetary damages defined for breach of contract.

It's an unenforceable clause. The affected party has no means to prove that a breach has happened.
I think we all ought to look at the ZDR fine-print here.

I get that in principle that there's no retention, but these are powerful models that can comprehend, paraphrase and summarize your logs for the sake of "product" improvement. Who knows what's collected here.

> It's an unenforceable clause. The affected party has no means to prove that a breach has happened.

Big tech spends hundreds of millions in high powered lawyers, audit logs, and contractual agreements with the sole purpose of proving your point wrong.

Anthropic (and probably OpenAI) can technically flag anything for 'safety' and view it, I'm not trying to imply that they are doing it arbitrarily, but they do have a way to do it within the ToS as far as I understand (happy to be proven wrong).
Are there terms clear? I dunno. There are ways to train on API usage without training on API usage.
How would you explain how they built https://github.com/anthropics/claude-for-legal then?

> ... for the legal workflows we see most

see.. where?

claude.com / chatgpt.com. By default your chats over there are opt-in for "help improve our models".

Not the API.

Terms for the app: https://privacy.claude.com/en/articles/10023580-is-my-data-u...

The API: https://privacy.claude.com/en/articles/7996868-is-my-data-us...

If you are a regular user they could care less,if you are enterprise they might be more careful. They have credit card info, and all your chats so I’m sure they can figure things out
Maybe Apple lawyers can figure this out in discovery.
If that was completely the whole story, then the extra Zero Retention etc tier enterprise things wouldn’t exist

I also don’t trust them lol

How much do we trust those guarantees? Apple currently alleges that OpenAI employees have been stealing Apple's trade secrets: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secre...

> [Mr. Tan] has directed job candidates still working for Apple to bring “Actual parts” from Apple to their interviews for “show and tell” sessions in which he and his team at OpenAI can elicit still more Apple confidential information.

> As part of its investigation, Apple found a “pattern by employees who depart for OpenAI of taking steps to evade the security processes intended to protect Apple’s confidential information.”

> Apple also claims former engineer Liu exploited a security bug to download confidential engineering files after leaving the company. Rather than report the exploit, Liu allegedly joked about it in messages (“LOL,” “so funny”). Liu also failed to return an Apple-issued laptop after his departure.

This seems pretty close to "they trust me, dumb fucks" behaviour.

They could use an agent to summarise the source material, and then train models on those summaries, and claim that some sort of clean-room training has happened?
You've been to too many meetings with PMs and directors saying "An agent could very easily do this"
I understand you're trying to be funny, but my point is that with novel technology there are novel ways to claim innocence in courts because of the legislative void.
my experience right now with fable via bedrock is they collect data, 5.6 has the option for ZDR at least
I think the risk of not doing is more existential than doing so and getting caught. Wouldn’t you agree?

Edit: And the point of the poster is they have already demonstrated a track record of lying and misconduct, so how can you trust their word now? What have they done to show you they have taken responsibility for past actions and changed?

lmao, wasn't xAI caught doing this recently? moreover at least moonshot is being honest about it.
they train on your requests by paraphrasing them (which means rewriting them but keeping all the saliency) and removing their association with you

i don't know why this is so controversial, their terms are written to perfectly fit this training regime. one of you downvoters i'm sure has an enterprise contract with them, just ask.

if you are using bedrock, until very recently, they didn't see your requests and could not paraphrase. but too many people were using bedrock for too much stuff they wanted to see. so that's why the terms for bedrock changed for fable 5. this was the core of the palantir / defense dept drama with anthropic.

Anthropic constantly uses dark patterns to steal training data from customers (like the “how is claude doing” spam, data retention loosening when the safeguards false positive, etc).
How is that a dark pattern? What is the light pattern for getting feedback from users?
There are multiple ways to use feedback. Personally I would often be fine with actual human reading my feedback, taking in account the points I made and evaluating how it should affect their feature development and future roadmap.

Now what I would expect AI companies to do is to take things which were submitted as feedback and pretty much adding to training:

"Do more of this: <copy of the whole response which was flagged as good in feedback>"

"Do less of this: <copy of the whole response which was flagged as bad in feedback>"

It's paraphrased, but the point is that they will most likely use it more-or-less as-is and thus whatever is in there will be part of the model's training set rather than someone picking up the parts from response that are important and only including them (which happens with traditional feedback).

If you have a naive loop like this you will quickly poison your dataset with e.g. thumbs downs because someone is unhappy with the latest frontend update, or teach models to not correctly refuse. I’m sure the pipeline is more sophisticated and in the middle.
But we already know they're not using the transcript for training (unless you opted in t that)
Not spam? There’s no “do not ask me again”.

Not typosquat? I responded with a sentence beginning in “1” once, and it jumped in during the race. It should have prompted with something like “WARNING: This will allow us to use this session including your source code for training, which is in violation of your account settings. Proceed with “Yes I understand”.

There is a world of difference between:

* A company following suit with their entire industry in choosing a very generous definition of fair use.

* A company being the first to defect and actually break their signed contracts with enormous enterprises committing to not train on those enterprises' most valuable assets.

Training on copyrighted works signs them up to be a part of a system that is at this point too big to fail and places them in good company with all of their competition. Breaking their signed agreements would open them up to very well-founded and well-funded lawsuits for contract violation and give their competition a huge boost.

All of a sudden "we actually don't break our contracts" would be a selling point. No company in their right mind is going to let what should be table stakes become a differentiator for their competition.

>> I pretty sure OpenAI and Anthropic are doing the same or worse.

So in your opinion, they are training on your data even if you toggle the "don't train on my data" checkbox off?

That's a bold assertion.

Not the guy you responded to, but I would assume ”they keep it safe” somewhere in a cold storage. Just in case they decide to train on it in a later phase.

Think of it as the Big Data hype some years ago.

I don't think they'd really be willing to risk the whole company on a small subset of prompts. It's not "keeping it safe", it's retaining proof of illegal activities.
Small subset of prompts? You mean literally every Kreti and Pleti who lets Claude go through their entire codebase is considered ”small subset of prompts”?

Kek

There is no evidence for these types of claims. They likely need to retain data for legal purposes (I think all of them are under injunctions from court cases), but there’s no way they will be breaching contracts with all these enterprises just for a little bit of data. Those contracts are their lifeline.
Yes, their entire existence relies on training on copyrighted content without permission being ok.
You truly see no difference between having a perhaps-overly-generous definition of fair use and flagrantly breaking contracts that you signed with your customers?
They both involve institutionalized lying.
Murder and speeding both involve violations of the law, there’s a big difference though.
Why wouldn't they?
Because the legal system does, in fact, have teeth. And those teeth actually deploy pretty readily. Especially when the people whose trade secrets you would be violating are gargantuan companies with enough resources that the cost of a lawsuit is a rounding error.
First it has to discover a violation.
Yeah but a disgruntled employee would talk sooner or later.
Obviously don't know for sure, but I can very easily seeing a combination of "move fast and break things", "it's easier to ask for forgiveness", "too big to fail", "I know tech, so I know everything", "AI is gonna change the world so fucking much, it doesn't matter what happens now", and finally "I cannot fail! I must make it work!" making especially con artist Sam just straight not care.
Pay to play teeth
To an extent, though for significant (in monetary terms) violations of the law the teeth tend to pay for themselves (but do so by not fully compensating the people whose behalf they are supposedly acting on).

More problematically there are camouflaged sharp spines pointed primarily in the direction of poorer people, and people not advised by lawyers.

But none of that matters here when the damaged parties include the megacorps of the world.

no it doesn't. If it would have teeth they would not resell copyright data. They will be busted like Kim DotCom
No AI company has been reselling copyright data to my knowledge, it would be truly bizarre if they did that.

What they have been doing, with some narrow exceptions where they have lost billions of dollars in court cases*, is not at all obviously prohibited by copyright law. Neither web scraping (i.e. asking for copies of data from people you have every reason to believe are authorized to give you copies) or running algorithms on copyrighted data are generally copyright infringment. I say generally because the "algorithm" of "ctrl-c ctrl-v" is obviously an exception, and there's some argument that training is similar enough to be illegal - a fairly weak argument that is mostly losing in court but has some tiny chance of still succeeding.

The law doesn't have teeth to prohibit things not prohibited under the law - no matter how much many people would like them to be prohibited. This shouldn't be surprising.

Unlike with copyright, the law does pretty clearly prohibit violating contractual terms to not hang onto or use other peoples data for purposes other than the narrow ones laid out in the contract when you agreed to the contract.

* Namely acquiring copies of data from people who they know aren't authorized to make copies - i.e. torrenting.

Does it? Because these companies systematically broke copyright law by illegally downloading terabytes of copyrighted content and there's been no consequences.

Past behaviour informs future trust and I wouldn't trust these companies whatsoever.

> and there's been no consequences.

Anthropic paid $1.5 billion for that, and never publicly deployed a model derived from the illegally downloaded data.

I'm not sure about the other companies off the top of my head - but I rather imagine they either never did this (I note that Google for instance already has lawfully acquired copies of basically every scrap of data you can imagine wanting to pirate) or are in the process of being sued or settled and I missed the news.

Because the value obtained from doing so is unlikely to exceed the cost of the lawsuits if they were ever caught doing so.
I work at OpenAI and I can assure you that if we say we don't train on your data, we don't.

I acknowledge that if you don't trust OpenAI, then you may not trust me either. But lying about this would be bad for legal liability, customer retention, and employee retention. Even if you model us as evil (and we really aren't), it's still not obvious to me that it would be a good decision to lie. As soon as a whistleblower revealed the scam, it would tank revenue and employee morale.

The reason why the models are getting better is training on users conversations.
[flagged]
I would also assume the same for non-Chinese as well
The nightmare for Anthropic to be caught doing that combined with the temptation of their staff to virtue-signal by blowing the whistle...

I trust them to act in their own interest if nothing else.

Aren't they on the "we, the safe AI, must win at any cost" justifiers?

What's a little contract violation if the fate of humanity is at stake?

And what if they use your data tò generate syntethic data to train on?
That would be just as bad.
I would assume at this point that any SaaS product, LLM or not, where the data doesn't reside on your servers on your premises (or your own colo) is training on the entire corpus of your data, whether they'll admit to it or not.
Not for Enterprise. You can safely assume the trillion dollar companies would ban GPT/Claude from being used in house if that was a concern.
Or think it the other way, what if both A and O actually train with these data, then the enterprises found it(or maybe never), what's the other options? I'm not trusting these LLM companies because a. their model live with human generated data even if they claim to generate data with their own model, it's just nit human data, b. this is business not charity, eventually they live with customers' data, no exception
We're talking about a dozen companies with trillion dollar market caps in the US that would each point their army of lawyers at OpenAI/Anthropic. Neither would survive that litigation. That risk far outweighs any potential gains in training data. Just not worth it.
I assume that all labs are training on any data they can get their hands on.
Yes, but open weights means you don't have to use their cloud providers. Someone in the US can host the model on their own infra.
And American providers, not sure if it's still the case but OpenAI were doing this.
I assume that of all of them as a basic security precaution.
This page says no, but the privacy policy is the authoritative document: https://www.kimi.com/help/kimi-api/api-data-security
You think openai, anthropic, google, z and any of the others dont? They do, if they say they dont, they do. Who wouldn't in this earth-shattering race. So Naive