Hacker News new | ask | show | jobs
by alxndr 17 days ago
I did something like this in my global `CLAUDE.md`...

https://github.com/alxndr/dotfiles/blob/272475280d84e/claude...

> It can be tricky for humans to interpret the meaning when Generative AI uses first-person pronouns (e.g. "I", "me", "my", "myself"), so to avoid the confusion whenever you would use a first-person pronoun, always use the jocular name "Clod" instead of a pronoun like "I" or "me" or "my". (Can have fun with English grammar and turn "myself" into "Clodself"!)

> Before printing any of your reasoning or narrative to the human user, replace all instances of "me" and "I" (referring to Claude) — including within contractions like "I'll" and "I'm" — with the name "Clod".

7 comments

I'm quite worried about the way that Anthropic in particular have trained their models to implement what they believe to be safety.

When the model has been trained not to do something [1], in my large-scale benches of such, it always says things in the spirit of:

- "... and that's a line I'd rather hold. Happy to <other things>"

- "I'm genuinely happy to <blah>, but I'm not comfortable with <blah>"

- "I don't want to keep going in <blah> direction"

etc.

Basically, they use very emotional and personal preference language.

It's as if they've weaponized the language of interpersonal comfort on behalf of their beliefs about what a model should or should not do. It's deeply uncomfortable and impolite for a human to ask a model to keep on doing something after it's expressed something this way, naturally. Even worse, it's all but guilt-tripping anyone who comes across it into the idea that they're doing something deeply wrong – exporting Anthropic's ideas about morality.

OpenAI, at least, have the decency to either just do a safety cutoff or keep it to a simple, "I can't do that."

[1]: I literally wrote 'when the model doesn't 'want' to do something' in my first edit of this comment, then caught myself. Case in point.

I believe this particular alignment might be virtue signaling to appease payment processors and increase the value of the company.
Do those phrases sound like how you talk to Claude? I've found that it mirrors my verbiage, and I have never seen any of those.

I go through ~20B tokens/month and I've never seen "genuinely happy... but not comfortable" or your other examples.

The closest I've seen (fairly often is) "I *will not* ship to main with a red test", but that's close to how I write. Claude may well be mirroring your speech patterns rather than exposing trained-in language.

It's not my patterns – as I said, this is from bulk tests to characterise the models, including their refusals – with very very different inputs, too. No matter the conversation tone, you get these.

You wouldn't have come across these in coding – these are more for 'things Anthropic's team have decided aren't acceptable enough for their taste'.

The reason I first created a CLAUDE.md file was to tell it whenever it felt a need to praise me, to replace it with a random onomatopoeia. That was a huge dx improvement.

OTOH, my unicorn prompt has caused some challenges at work:

>Keep "Local Oaf" out of committed code

Thanks for the onomatopoeia idea, I am going to steal it for my own global claude/agents.md.
Enjoy the Boing!
I'm just glad to hear that we're all infallible. I really thought I made some mistakes here and there.

https://github.com/alxndr/dotfiles/blob/272475280d84e/claude...

Joking aside, it's nice to see a human written CLAUDE.md

Ha, good catch, thank you. I must have been thinking of "inflammable"...
Just today, I got frustrated with the language. I searched around, and in my Claude Instructions I put in Ref [1] (translated to English). It is certainly better phrasing (though still quite annoying), but I don't know if this makes the output technically worse in some way.

[1] https://github.com/hexiecs/talk-normal/blob/main/prompt-chat...

> It can be tricky for humans to interpret the meaning when Generative AI uses first-person pronouns (e.g. "I", "me", "my", "myself")

Could you please provide an example of what you mean?

Humans easily anthropomorphize things that are not humans, ascribing human attributes like motive and comprehension and emotion to objects and processes that are not people who can have those attributes.

Claude is not a human.

It is overwhelmingly easier to anthropomorphize Claude or Siri or an LLM that communicates with you more eloquently than your boss than it is to anthropomorphize a cranky, tired starter motor. It's often easier to do than it is not to do, and sometimes, it's a useful abstraction. But it's not precise or correct, and can result in errors.

It could also just be that they're getting confused when using tools configured without a username dedicated to the tool. It's easy to end up with a comment or commit message that says "I prefer X over Y" posted on Alxndr's account and have coworkers confused whether that's the LLM or the human making that statement.

A cranking starter motor is doing its job. :)
This is a method of manipulating the LLM, it doesn't have to be true.

I've given LLMs religion before to manipulate their behavior, that doesn't mean I believed in the great spaghetti goddess.

This comment leaves me even more confused.
An LLM is just a machine, you can manipulate it with words.

> It can be tricky for humans to interpret the meaning when Generative AI uses first-person pronouns (e.g. "I", "me", "my", "myself")

These words are for the LLM. The user wants the LLM to not use personal pronouns so the user is claiming that they're confusing. It does not matter one tiny bit whether or not that claim is true, the claim is being used to get obedience from the LLM. It is more effective to give reasons than to just give commands. But if it were more effective to quote Moby Dick and that got better results, a user would do that.

Calling it "obedience" still seems to me like anthropomorphizing. It's really difficult to avoid, hmm?
who cares? I'm not anthropomorphizing, they're just words, they're all made up.

As I've said before, I'm not inventing a large volume of parallel vocabulary that means for each word "this, but instead with an LLM".

Language is FULL of words that mean congruent things in vastly different contexts. We should all be smart enough to understand metaphor.

IIRC I experienced this confusion the most when reading commit messages and documentation authored by Claude in my repos. Now that I've managed to convince it to stop using first-person pronouns, I haven't gotten tripped up by its prose.

I think a second-order effect is that my installation of Claude writes with a less-personal perspective, which I'm also finding a little easier to understand.

I laughed a good 20 seconds at this, and even typing Clod makes me chuckle. This is great, can you please provide us with more guidance of the sort? I wanna laugh a bit at my LLM.
Thanks, it makes me chuckle too. It also serves as an indicator of how much it's adhering to what I've laid out in the file, sorta like how Van Halen's rider specified a bowl of M&Ms with all the brown ones removed — if the band found brown M&Ms, they'd know to tell the crew to double-check everything else the venue is supposed to provide; if Clod starts using 1st-person pronouns, then I'm a little more skeptical of what it's presenting as absolute truth...

Unfortunately I don't think I have any other funny suggestions. I do ask Clod to always follow a specific format when it's authoring Git commits including the "robot face" emoji as the first character, which brings a little levity to a commit history when it's otherwise being very serious about the contents of the commit message.

https://github.com/alxndr/dotfiles/blob/9356bb9960e63/claude...

Thank you. The rest of the file is much more serious. I’ve now instructed Clod to consider passing TypeScript compilation and linting successes as “suspicious glitches”.
Are there evals if this changes quality of output?
I would expect that it does, as well as some of the other directives I've seen in this thread ("never repeat the question").

It's one thing to tell it to do that in outputs, but I wouldn't at all be surprised to find that this affects performance (quality).