> "I want LLMs to code for me, but I want them to be trained on other people's code, not mine, duh".
Who ever said that? Have you actually heard that from your fellow programmers in real life?
If the code I wrote actually made even the slightest discernible difference in LLMs I'd be so honored. But it won't happen, as it's just 0.00001% of all the training data.
Tragedy of the Commons is not an analogy, it is an inevitable result of a large fraction of participants in a coordination game defecting due to perceived individual advantage.
Sounds good? They can pay for code they want to train on. There are plenty of companies sending me offers to code training materials for them for $50-100/hr. Don’t expect to charge me an arm and a leg for inference and then also train on my code.
There are already opt out buttons for training in Cursor and Claude Code… if you don’t want it then turn it off. If it was worth enough money to them they would offer a monetary incentive like discounts but none of them have yet
Interesting how our generation which grew up using Napster now has so many intellectual property extremists. By this logic, even humming a tune you heard on the radio is theft.
> I have absolutely zero interest in free. I honestly don't think I'm even remotely in the same demographic as people using free tiers / models. I want to pay. I don't want my data used for training...
They want to use LLMs trained on others code but don't want to contribute with their own.
It makes sense from a business perspective-SaaS firms value the ability of coding agents to accelerate development, but also worry the models will learn the secret sauce of their business and destroy its moat. So their desire to contractually exclude training on their data has some logic to it.
(Disclaimer: Not speaking for or about my current employer, just a general industry observation.)
I don't really use LLMs myself, but if someone wants to have any kind of software business then having the models trained on their products isn't ideal.
Who ever said that? Have you actually heard that from your fellow programmers in real life?
If the code I wrote actually made even the slightest discernible difference in LLMs I'd be so honored. But it won't happen, as it's just 0.00001% of all the training data.