Hacker News new | ask | show | jobs
by sanxiyn 1 day ago
Any reasonable safety testing should include finetuning and safety margin to account for others may do better finetuning.
1 comments

I can fine tune significant behavior changes, there is little model developers can do to prevent this (aiui), so this effectively becomes an blanket ban
Yes, I agree it is effectively a blanket ban (above some capability) for now. I hope AI alignment research advances in the future so that it is not so.
a ban is effectively impossible without a global treaty

the current US admin as pulled out and worked against all sorts of global treaties, agreements, and negotiations; sending the president's friends instead of experts; who's going to trust us?