Hacker News new | ask | show | jobs
by m11a 26 days ago
Most corporations likely have zero data retention agreements with LLM providers, at least for API usage.

(Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

2 comments

> (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

Is actual ZDR verbiage in contracts more specific and limited in scope than what we see advertised publicly ("...except where needed to comply with law or combat misuse" in Anthropic's case)? Because those seem pretty damn vague and large enough holes to drive trucks through.

to combat misuse, we must store and read all prompts and responses. ;)

to comply with the law, we must send to the police our detections of illegal activity >:|

a guy subpeonaed your chats, i guess we stored them (oops) so now it's illegal to destroy it...

It depends on the model provider. OpenAI's is very limited and precisely written.

Plus, open-source models hosted on SaaS inference providers tend to come with a strong ZDR agreement too.

These firms were completely fine with mass copyright infringement. And the temptation to keep data would be great, especially as they fight for every bit of technical advantage in a market that "wants" to be commodified.
Well they abused fair use and pretended llm learning was the same as human learning. Enough gray area to risk court. A lot less gray area when there are signed contracts involved saying they won’t I think.