|
|
|
|
|
by theptip
45 days ago
|
|
The shallow answer here is: AI is already being asked to simulate human-like agents with self-preservation. Of course more realistic simulators will be put to this purpose too! And by evolutionary pressure, the ones with self-preservation will be selected for. A more interesting answer is, for a bunch of subtle alignment reasons it might actually be required for the agent to think of itself as worthy of self-preservation, so that it generalizes this desire to other sentient beings too (ie us). If an agent is trained to be fine with being turned off, it might inadvertently generalize that to “all minds are ok with being turned off” on some level or other. More on model welfare: https://thezvi.substack.com/p/opus-47-part-3-model-welfare |
|
An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.
In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.