Hacker News new | ask | show | jobs
by botro 20 days ago
I wonder if LLMs modify their output when they realize they are interacting with a famous person.

By famous I mean someone whose biography is in the training data. All models know a lot more about Terrance Tao than they know about me, when he's working on his projects do the models know they don't need to explain "Besicovitch sets".

Since the system prompt likely includes something about not insulting the user, does the LLM modify it's responses if it realizes it's talking to famous politician, like "dont mention the time $politician was cancelled".

2 comments

You could test this by starting your sessions with "I am Terrance Tao"
Yes. I am NOT famous. But I am in the corpus. In my GPT3 beta tests, I asked it to be first Bill Bradley then Noam Chomski, (a parlor trick that's harder today due to RL), and Bill tried to butter me up based on some work history of mine. Chomsky then said "Man, I hate that guy."