Hacker News new | ask | show | jobs
by pietz 12 days ago
[flagged]
5 comments

I would also assume the same for non-Chinese as well
The nightmare for Anthropic to be caught doing that combined with the temptation of their staff to virtue-signal by blowing the whistle...

I trust them to act in their own interest if nothing else.

Aren't they on the "we, the safe AI, must win at any cost" justifiers?

What's a little contract violation if the fate of humanity is at stake?

And what if they use your data tò generate syntethic data to train on?
That would be just as bad.
I would assume at this point that any SaaS product, LLM or not, where the data doesn't reside on your servers on your premises (or your own colo) is training on the entire corpus of your data, whether they'll admit to it or not.
Not for Enterprise. You can safely assume the trillion dollar companies would ban GPT/Claude from being used in house if that was a concern.
Or think it the other way, what if both A and O actually train with these data, then the enterprises found it(or maybe never), what's the other options? I'm not trusting these LLM companies because a. their model live with human generated data even if they claim to generate data with their own model, it's just nit human data, b. this is business not charity, eventually they live with customers' data, no exception
We're talking about a dozen companies with trillion dollar market caps in the US that would each point their army of lawyers at OpenAI/Anthropic. Neither would survive that litigation. That risk far outweighs any potential gains in training data. Just not worth it.
I assume that all labs are training on any data they can get their hands on.
Yes, but open weights means you don't have to use their cloud providers. Someone in the US can host the model on their own infra.
And American providers, not sure if it's still the case but OpenAI were doing this.
I assume that of all of them as a basic security precaution.