|
|
|
|
|
by dmrivers
1 day ago
|
|
I wonder how much of the socially left results are affected by the "harmlessness" part of the RLHF post-training? Companies don't want to be sued over LLMs that recommend harm in any way, so RLHF pushes them to say no to "death penalty", "spanking", and "incarceration" which are all violence-coded. A lot of this test seems to be about willingness to be violent, which LLMs are generally unwilling to be. |
|