|
|
|
|
|
by notahacker
14 days ago
|
|
Feels like both the sycophancy and the nitpicking come from it being a RLHF-ed probabilistic model of text continuations, in which generating a continuation which politely pointing out common minor errors is generally treated as an optimal response, and pointing out that a subtle flaw makes the whole exercise futile isn't necessarily, even if the agent has the ability to identify that flaw and extrapolate its consequences for the whole project which a cheap or instant model likely doesn't. |
|