|
|
|
|
|
by demosthanos
14 days ago
|
|
Claude's "honest" is an interesting example because we can trace it to a specific document that it was trained on extensively: the "Constitution" is identified to Claude in its training as the core of what it is, and it uses the word "honest" or a derivative 57 times, including having a whole section on it. > Honesty is a core aspect of our vision for Claude’s ethical character. Indeed, while we want Claude’s honesty to be tactful, graceful, and infused with deep care for the interests of all stakeholders, we also want Claude to hold standards of honesty that are substantially higher than the ones at stake in many standard visions of human ethics. https://www.anthropic.com/constitution |
|
But Sol actually has the same obsession with honesty: I suspect it's more an artifact of trying to control reward hacking.
Models will lie, obfuscate, and mislead under the pressure of RL, so both OAI and Ant are probably forced to spend a lot of time coaxing "honest" answers out of the model
OpenAI's recent prompt for a math conjecture hints at a lot of it when instructing on subagents: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...