|
|
|
|
|
by SonOfLilit
26 days ago
|
|
The "OG" alignment research that MIRI were publishing long before LLMs burst into the scene spent most of it's time on that question. "How can we even define what an aligned AI should do, if human's are not aligned with each other?" as well as "What does being aligned mean when you're a wizard box who's main influence on the world is to create stronger wizard boxes?" and other deep philosophical questions. They came up with a framework called Coherent Extrapolated Volition to address this specific question. https://en.wikipedia.org/wiki/Coherent_extrapolated_volition |
|
First of all, calling it “coherent” extrapolated volition presupposes that there is such a thing. It doesn’t actually address the objection above, that there may be no such thing. It’s a bit like saying you solved car safety by presupposing a safe car.
Second, it assumes that such a thing can be effectively measured, and there will be no problems or controversies with the extrapolation process itself. There may be several EVs to choose from, and at that point the framework has nothing to say. Maybe we just pick at random then I suppose.