These firms were completely fine with mass copyright infringement. And the temptation to keep data would be great, especially as they fight for every bit of technical advantage in a market that "wants" to be commodified.
Well they abused fair use and pretended llm learning was the same as human learning. Enough gray area to risk court. A lot less gray area when there are signed contracts involved saying they won’t I think.