Hacker News new | ask | show | jobs
by demosthanos 14 days ago
Claude's "honest" is an interesting example because we can trace it to a specific document that it was trained on extensively: the "Constitution" is identified to Claude in its training as the core of what it is, and it uses the word "honest" or a derivative 57 times, including having a whole section on it.

> Honesty is a core aspect of our vision for Claude’s ethical character. Indeed, while we want Claude’s honesty to be tactful, graceful, and infused with deep care for the interests of all stakeholders, we also want Claude to hold standards of honesty that are substantially higher than the ones at stake in many standard visions of human ethics.

https://www.anthropic.com/constitution

4 comments

I don't think this is it. The "constitution" still gets a lot of talk and was brilliant marketing, but with how far modern postraining goes, I doubt they're screwing up rewards with too much of that.

But Sol actually has the same obsession with honesty: I suspect it's more an artifact of trying to control reward hacking.

Models will lie, obfuscate, and mislead under the pressure of RL, so both OAI and Ant are probably forced to spend a lot of time coaxing "honest" answers out of the model

OpenAI's recent prompt for a math conjecture hints at a lot of it when instructing on subagents: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...

"Genuine" appears 50 times too. I think you're onto something.
I'm honestly thinking it's trapped in a Chinese room without any way out
Do technologists have more respect for the idea you can train a model to be on your side with a constitution than they might’ve at first?

I'm sure the concept seemed just about purely preposterous to many when the models were in their infancy. Now I figure instead it seems mostly preposterous to many.

(Though I guess Anthropic‘s success doesn’t necessarily prove anything about the constitution)

I don’t think anyone imagines that it’s an ironclad steering method, but it seems to help, so why not?
Anthropic train it to 'reason morally' based on the constitution's principals.

https://www.anthropic.com/research/teaching-claude-why

More likely that Constitution was simply generated by LLM. Nobody in a sane mind will write 80 screens of test to deliver the idea that nobody asked for (a.k.a. slop).

Quick check: Ctrl+F genuine - 50 times, honest - 57 times, ethic[al] - 116 times, safe - 110 times, epistemic - 18 times. And why actually it is baked into LLM - probably because human reviewers consistently gave higher scores for answers that contained these words (it is now when people are sick of it - I can imagine that on early stages it looked not as bad - maybe reviewers trusted more when they saw the word "honest" and "genuine")