|
|
|
|
|
by wyrdcurt
8 days ago
|
|
In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models. |
|