|
|
|
|
|
by teravor
4 days ago
|
|
it's not actually possible to make an open model that cannot be easily jailbroken. when you have access to its entire state it's trivial to gaslight it into a non-refusal state (you can forge its responses to build up the jailbroken state). I don't know if the sole kimi k3 provider gives that level of access atm however (where you can dictate its own responses to it). |
|
Agreed, but you can make a model that doesn't know much about a topic. gpt-oss is pretty well known for not being trained on erotica stuff. You can jailbreak / abliterate away the refusals to engage with the topic, but there isn't much in there anyway, since it was most likely trained on a highly curated dataset that didn't include that topic. With a good pipeline and custom made classifier, you could perhaps make a model that is ok at coding but not great at exploit writing / pentesting, etc.