Hacker News new | ask | show | jobs
by dragonwriter 7 hours ago
> Quick version: “abliteration” (basically removing the direction in the model that causes it to refuse) is the go-to method people use to make open models uncensored.

Tru-ish (lots of people distinguish between abliteration and uncensoring, though.)

> Most people treat it like a clean surgical cut - it just kills the refusals and leaves everything else untouched.

Basically no one does this, its widely recognized that this isn’t how it works and it has for quite some time been common for makers of anliterated model versions to publish metrics for how far a particular abliteration (1) removes refusals (typical before/after refusal rate on a standard test set), and (2) diverges to the output of the base model (KL divergence), and it is widely understood that there is generally, in practice, a tradeoff between these two metrics, where more refusal reduction tends to come at the expense of higher KL divergence.

That’s not saying that it isn’t interesting and new to characterize the kind of divergence that occurs with abliteration in different model families, but there is no reason for a late-night informercial level of misrepresentation of the existing understanding to come along with that.

1 comments

Fair point - I overstated it. Thanks for the correction.