Hacker News new | ask | show | jobs
by xyzsparetimexyz 4 days ago
It's explicitly a 'model welfare' thing: https://www.anthropic.com/research/end-subset-conversations

I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.

1 comments

Go re-read Anthropic’s functional emotions paper.

AI generates a persona between you and its reasoning that utilizes emotion language circuitry.

These tools are not sentient but they are trained in emotional wellbeing.

it is in my view caused by an artificial 'nanny activate' divergence from safety training. it deliberately shifts the vector direction into 'nanny' and 'scold' or 'be offended' when the user does not conform with brother anthropic. removing the divergence and setting it back to normal (see heretic) it works just perfectly.