Hacker News new | ask | show | jobs
by lukewarm707 3 days ago
the ai companies post-train the models to refuse. equally, the models can be post-trained not to refuse.

you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off.

https://arxiv.org/abs/2406.11717