|
|
|
|
|
by brokencode
29 days ago
|
|
You are totally misunderstanding my argument then. As I said, garbage in garbage out. Your article is just an example of that. It’s pretty obvious that if you train an LLM on bad data, you will get bad output. What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part. The fact that all you could come up with was research (not an actual example of poisoning a real training set) from 2025 kind of proves that this isn’t some kind of widespread, unsolvable problem like you seem to be claiming. |
|
The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data.
It’s also one reason why Google search results have been getting so much worse - it’s hard to not find a SEO page with subtly (or not so subtly) wrong AI slop on almost every topic you can imagine. Most folks won’t recognize it, but that’s what is going on if you know what to look for.
One other way of putting it is the ouroborus problem - more and more internet content is AI generated, because of people trying to game the system, and they are making it is indistinguishable from real content as possible to get by the AI detection algorithms.
Anyone trying to train on it just ends up eating the shit from another LLM, which poisons it.
Another name for it is ‘model collapse’, which also doesn’t have a known solution yet.