Hacker News new | ask | show | jobs
by Melonai 19 days ago
I noticed that too, I guess the content in the training data suggests that when some text self-identifies as "cold, hard logic" it additionally assumes a mean tone. Additionally it makes responses more likely to disagree even when given inputs that are more or less valid. A while ago I tested out giving some idea to a model with the "hard logic" prompt, then taking its own output, reverting the conversation and giving it its own statement again, making it disagree with the very same statement it just made itself. This was mostly just humorous, and doesn't indicate too much, after all both statements might've been wrong in some way, but the statement itself was moreso an opinion with no fully right answer, showing this tendency to disagree.

I guess when we tell models to use "cold logic" they don't interpret that purely definitionally, but moreso with what this statement usually implies, which is often a disagreement between two people, one trying to leverage supposed logic to disagree and denigrate the other party, oftentimes actually arguing out of emotion and not the logic they claim to use. This probably occurs enough to give model responses a mean tint. That's my theory on why this could occur at least.