|
|
|
|
|
by overgard
1 day ago
|
|
The notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth. But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable! In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it. IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else. I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities. This is all kayfabe and marketting. |
|
Is it not also one of the most important use cases for AI to apply existing knowledge to new applications?
As a hopefully exaggerated example, I would think one could apply knowledge about pesticides, chemistry, and medicine to create biological weapons.