|
|
|
|
|
by eddyg
4 days ago
|
|
Constitutional classifiers go a long way to reducing unsafe usage in closed-weight models. And like we saw with Fable, closed models can be revoked and classifiers updated when “jailbreaks” are found. Having the weights gives you the exact affordance an unlearning attack requires, without rate limits. |
|