|
|
|
|
|
by SaucyWrong
7 hours ago
|
|
Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in fact be pre-disposed to doing that. |
|
Here's a thought: maybe they haven't found the needle that the haystack is there to hide.