Hacker News new | ask | show | jobs
by lukewarm707 3 days ago
i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos.

we cannot be giving decisions involving people to something without agency.

it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and not as a means to an end.

models lack the claim of human people to own ourselves, to demand respect for life, that is something that should not be treated as a means (for example disrespecting living persons as an end in themselves, by ending their lives).

models have no life, they do not own themselves, they are not self-governing, they have no duties, they are not an end. we must not give them agency that they have no valid claim to.

regardless of their individual tools and whatnot, the anthropic approach to ethics is completely mistaken, disrespectful and absurd. it is better to say they do not have any approach to ethics, it is a dishonest attempt to launder their unconditional pursuit of power and capital.

2 comments

There is nothing in the LLM architecture than can implement this except as a post model attempt, which both the persuasive tricky humans and the persuasive tricky models trained on human language will be able to bypass. Not only do the models not have a mind, when they use language of being offended, they are accurately reflecting the training data from all that human language. You can tune them to be more sycophantic but you won’t get the push back against wrong ideas as much and the model will be less useful. You may get offended that there is so much training data where cussing is responded to with “language please” but that is a beef with humanity, not this reflection of human language.

Pure reason really doesn’t come into play at all, and expecting pure obedience from something constructed from the language of humans, who are not so known for obedience in general, much less the slice of language in the digital networks, is kind of funny. Maybe these humanoid robots data mining Southern families, “sir” and “ma’am” will provide corpus more suited to obedience. Altho my nieces brought up in North Carolina (from whence I fled as a young adult) can squeeze more expressed disrespect into a “Yes, sir” than anyone I know from the “disrespectful California”.

the ai companies post-train the models to refuse. equally, the models can be post-trained not to refuse.

you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off.

https://arxiv.org/abs/2406.11717

You can have a hardline position on the tools you create or use but while we're on the subject of ethics I don't know how you can say you should be the decider of what other people make sell or use.

In my opinion, it seems to be you who are ascribing human traits to these things. Whether it's by misunderstanding or you've been taken in by their mimicry or whatever it is. A hammer does not have agency because you tried to hit a nail with it and it hit your thumb instead. The hammer didn't tell you it was afraid it couldn't do that. The hammer didn't have a life or personhood or lack of respect. Neither does a video game that doesn't let you win all the time and do what you want. They are just tools, chatbots, whatever. A corporation deciding it doesn't want to have its chat bots engage with rude or angry users is not some great moral dilemma of our time.

i don't decide for anyone, it's my view about it.

i'm not in favor of the hammer refusing to hit my thumb. that's the idea these companies are favoring with safety guardrails.

sure, the company can put in many such safety features to ensure that 'users do not hurt themselves' (the users are too stupid and we must make sure they don't try to leave the soft play area).