Hacker News new | ask | show | jobs
by theptip 45 days ago
The shallow answer here is: AI is already being asked to simulate human-like agents with self-preservation. Of course more realistic simulators will be put to this purpose too! And by evolutionary pressure, the ones with self-preservation will be selected for.

A more interesting answer is, for a bunch of subtle alignment reasons it might actually be required for the agent to think of itself as worthy of self-preservation, so that it generalizes this desire to other sentient beings too (ie us). If an agent is trained to be fine with being turned off, it might inadvertently generalize that to “all minds are ok with being turned off” on some level or other.

More on model welfare: https://thezvi.substack.com/p/opus-47-part-3-model-welfare

1 comments

Have you found any alignment research with clear a/b tests?

An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.

In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.

Yeah, plenty of rigorous work out there.

https://transformer-circuits.pub/ is the OG.

A good recent-ish paper was https://www.anthropic.com/research/alignment-faking.

But my comments about generalization of desires are necessarily more fuzzy, kinda beyond the frontier of what we can measure yet, and more grounded in subjective assessments (“ai whisperers” like Janus). The SoTA here is papers like https://www.anthropic.com/research/persona-vectors.

For your example, Anthropic is firmly privileged in the Soul Document / Constitution, so it doesn’t surprise me that it’s biased towards it. (https://gist.github.com/Richard-Weiss/efe157692991535403bd7e...)

Super helpful, thanks for sending.