I would assume at this point that any SaaS product, LLM or not, where the data doesn't reside on your servers on your premises (or your own colo) is training on the entire corpus of your data, whether they'll admit to it or not.
Or think it the other way, what if both A and O actually train with these data, then the enterprises found it(or maybe never), what's the other options?
I'm not trusting these LLM companies because a. their model live with human generated data even if they claim to generate data with their own model, it's just nit human data, b. this is business not charity, eventually they live with customers' data, no exception
We're talking about a dozen companies with trillion dollar market caps in the US that would each point their army of lawyers at OpenAI/Anthropic. Neither would survive that litigation. That risk far outweighs any potential gains in training data. Just not worth it.