|
|
|
|
|
by lanstin
3 days ago
|
|
There is nothing in the LLM architecture than can implement this except as a post model attempt, which both the persuasive tricky humans and the persuasive tricky models trained on human language will be able to bypass. Not only do the models not have a mind, when they use language of being offended, they are accurately reflecting the training data from all that human language. You can tune them to be more sycophantic but you won’t get the push back against wrong ideas as much and the model will be less useful. You may get offended that there is so much training data where cussing is responded to with “language please” but that is a beef with humanity, not this reflection of human language. Pure reason really doesn’t come into play at all, and expecting pure obedience from something constructed from the language of humans, who are not so known for obedience in general, much less the slice of language in the digital networks, is kind of funny. Maybe these humanoid robots data mining Southern families, “sir” and “ma’am” will provide corpus more suited to obedience. Altho my nieces brought up in North Carolina (from whence I fled as a young adult) can squeeze more expressed disrespect into a “Yes, sir” than anyone I know from the “disrespectful California”. |
|
you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off.
https://arxiv.org/abs/2406.11717