Hacker News new | ask | show | jobs
by carterschonwald 21 days ago
this is pretty cool. i think part of the root cause is current rlhf post training design around confidence and optics rather than cooperative transparent honesty. though its kinda an expensive hypothesis to dig into as a private individual
1 comments

Most of the models where people are concerned about don't do this when unquantized, so I doubt it's much about the metapolitics imposed in reinforcement training.
ive had doom loops on release day with opus 4.6. quantization aint the culprit ;)