|
|
|
|
|
by Salgat
33 days ago
|
|
This is a misleading statement. The "private data" is still largely publicly produced data that has been curated through private agreements instead of scraping, such as reddit posts/comments (this is the "third-party data agreements" that companies like OpenAI mention). And yes, there is still a lot of processing done on this data, which is the norm for preparing training data. |
|
Source: Work at a lab, common knowledge.