|
|
|
|
|
by poisonfountain
16 days ago
|
|
My theory is that the labs are RL'ing models to output easily classifiable text so they can avoid model collapse when training the models on data scraped from the web. Of course, you can skew the distribution with some effort and generate text that avoids even the best classifiers out there (like Pangram), but even tech-savvy people aren't usually doing it (see the amount of AI-written posts that end up in HN and get tons of comments complaining about AI mannerisms), so I guess they're successfully avoiding like 99% of the slop using such classifiers. I don't think it's in the interest of the labs to allow you to generate text that's indistinguishable from human prose. Especially since nobody would pay $1,000/mo just to generate text - but would do so for tasks like coding. |
|