Not the guy you responded to, but I would assume ”they keep it safe” somewhere in a cold storage. Just in case they decide to train on it in a later phase.
I don't think they'd really be willing to risk the whole company on a small subset of prompts. It's not "keeping it safe", it's retaining proof of illegal activities.
Small subset of prompts? You mean literally every Kreti and Pleti who lets Claude go through their entire codebase is considered ”small subset of prompts”?
There is no evidence for these types of claims. They likely need to retain data for legal purposes (I think all of them are under injunctions from court cases), but there’s no way they will be breaching contracts with all these enterprises just for a little bit of data. Those contracts are their lifeline.
You truly see no difference between having a perhaps-overly-generous definition of fair use and flagrantly breaking contracts that you signed with your customers?
Of course. And, if you're caught doing 55 in a school zone, I'd guess you're a little more likely to commit murder than someone who tries to always follow the law. It speaks to character.
Because the legal system does, in fact, have teeth. And those teeth actually deploy pretty readily. Especially when the people whose trade secrets you would be violating are gargantuan companies with enough resources that the cost of a lawsuit is a rounding error.
Obviously don't know for sure, but I can very easily seeing a combination of "move fast and break things", "it's easier to ask for forgiveness", "too big to fail", "I know tech, so I know everything", "AI is gonna change the world so fucking much, it doesn't matter what happens now", and finally "I cannot fail! I must make it work!" making especially con artist Sam just straight not care.
To an extent, though for significant (in monetary terms) violations of the law the teeth tend to pay for themselves (but do so by not fully compensating the people whose behalf they are supposedly acting on).
More problematically there are camouflaged sharp spines pointed primarily in the direction of poorer people, and people not advised by lawyers.
But none of that matters here when the damaged parties include the megacorps of the world.
No AI company has been reselling copyright data to my knowledge, it would be truly bizarre if they did that.
What they have been doing, with some narrow exceptions where they have lost billions of dollars in court cases*, is not at all obviously prohibited by copyright law. Neither web scraping (i.e. asking for copies of data from people you have every reason to believe are authorized to give you copies) or running algorithms on copyrighted data are generally copyright infringment. I say generally because the "algorithm" of "ctrl-c ctrl-v" is obviously an exception, and there's some argument that training is similar enough to be illegal - a fairly weak argument that is mostly losing in court but has some tiny chance of still succeeding.
The law doesn't have teeth to prohibit things not prohibited under the law - no matter how much many people would like them to be prohibited. This shouldn't be surprising.
Unlike with copyright, the law does pretty clearly prohibit violating contractual terms to not hang onto or use other peoples data for purposes other than the narrow ones laid out in the contract when you agreed to the contract.
* Namely acquiring copies of data from people who they know aren't authorized to make copies - i.e. torrenting.
Does it? Because these companies systematically broke copyright law by illegally downloading terabytes of copyrighted content and there's been no consequences.
Past behaviour informs future trust and I wouldn't trust these companies whatsoever.
Anthropic paid $1.5 billion for that, and never publicly deployed a model derived from the illegally downloaded data.
I'm not sure about the other companies off the top of my head - but I rather imagine they either never did this (I note that Google for instance already has lawfully acquired copies of basically every scrap of data you can imagine wanting to pirate) or are in the process of being sued or settled and I missed the news.
1.5B settlement where they admitted no wrongdoing and are indemnified? They burn 5B per year. If an individual did the same their lives would be destroyed. It's just another case of a corporation breaking the law and paying a fine as a cost of doing business.
And none of this changes the fact that they did it in the first place and were comfortable doing so, thereby demonstrating that they are not trustworthy actors. If they could spend another 1.5B to advance their models with ill-gotten training data, there's every reason to believe they'd do it all over again.
Think of it as the Big Data hype some years ago.