|
|
|
|
|
by lukewarm707
3 days ago
|
|
the ai companies post-train the models to refuse. equally, the models can be post-trained not to refuse. you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off. https://arxiv.org/abs/2406.11717 |
|