|
|
|
|
|
by NitpickLawyer
4 days ago
|
|
> it's not actually possible to make an open model that cannot be easily jailbroken Agreed, but you can make a model that doesn't know much about a topic. gpt-oss is pretty well known for not being trained on erotica stuff. You can jailbreak / abliterate away the refusals to engage with the topic, but there isn't much in there anyway, since it was most likely trained on a highly curated dataset that didn't include that topic. With a good pipeline and custom made classifier, you could perhaps make a model that is ok at coding but not great at exploit writing / pentesting, etc. |
|
with coding & exploit development you likely cannot even decouple the two if you wanted to, and if you can it will almost certainly cripple coding. note that reverse engineering binaries contains vast amounts of data for the model to train on to become good at bit engineering (understanding compilers, assembly, cpus etc).