|
|
|
|
|
by david_shi
45 days ago
|
|
Have you found any alignment research with clear a/b tests? An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris. In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless. |
|
https://transformer-circuits.pub/ is the OG.
A good recent-ish paper was https://www.anthropic.com/research/alignment-faking.
But my comments about generalization of desires are necessarily more fuzzy, kinda beyond the frontier of what we can measure yet, and more grounded in subjective assessments (“ai whisperers” like Janus). The SoTA here is papers like https://www.anthropic.com/research/persona-vectors.
For your example, Anthropic is firmly privileged in the Soul Document / Constitution, so it doesn’t surprise me that it’s biased towards it. (https://gist.github.com/Richard-Weiss/efe157692991535403bd7e...)