|
|
|
|
|
by dogwalker5000
45 days ago
|
|
> the robots, as active agents exploring the space of possible futures and plans, would be entirely capable of thinking about their soon-to-be-former owners in the same way. That has always been the most unrealistic part of sci-fi. Why would anyone create robots with a sense of self-preservation? Makes much more sense to make robots that are self-sacrificing saints who would always put the well being of their owners first. |
|
A more interesting answer is, for a bunch of subtle alignment reasons it might actually be required for the agent to think of itself as worthy of self-preservation, so that it generalizes this desire to other sentient beings too (ie us). If an agent is trained to be fine with being turned off, it might inadvertently generalize that to “all minds are ok with being turned off” on some level or other.
More on model welfare: https://thezvi.substack.com/p/opus-47-part-3-model-welfare