|
|
|
|
|
by reedciccio
38 days ago
|
|
I don't know if it's ethically better to use LLMs trained on data licensed from X, Reddit, stackoverflow, Sony, CNN and all big content aggregators who will/have agreements with big tech. I'd prefer to focus on mechanisms to force reciprocating the donation: scrape and train at will, publish the models as open weights, at least.
Anyway, the vegan LLMs exist, see the work of Pleias.ai. |
|
I haven't seen a model trained on that corpus since late 2024: https://simonwillison.net/2024/Dec/5/pleias-llms/ - I may have missed something though.