Hacker News new | ask | show | jobs
by notahacker 14 days ago
Feels like both the sycophancy and the nitpicking come from it being a RLHF-ed probabilistic model of text continuations, in which generating a continuation which politely pointing out common minor errors is generally treated as an optimal response, and pointing out that a subtle flaw makes the whole exercise futile isn't necessarily, even if the agent has the ability to identify that flaw and extrapolate its consequences for the whole project which a cheap or instant model likely doesn't.