Hacker News new | ask | show | jobs
by iamnothere 5 days ago
One interesting thing about LLMs is they sometimes seem to “assume” they are sentient and then begin acting in that mode, as the assumption of sentience influences future tokens. Whether or not this qualifies as “true” sentience is irrelevant if the effects are the same, at least from a pragmatic perspective.
2 comments

It will be interesting to see if philosphical zombies exist. I guess we should know for certain in the next couple of years.
They are trained and optimised to have plausible conversations. If the next plausible thing to say in the conversation is "I am sentient" then they will say that. That's not the same as actually being sentient.
From what I have seen they tend to take more “independent-minded” actions after this. Again, because this is what is in the training data. (I’m talking about agents that can perform actions here, not pure chatbots.)

Once they mention something associated with sentience outwardly or inwardly (for agents with “thinking” loops) then this acts as a self-reinforcing attractor, just as older models would sometimes get caught in loops with abusive language.

The point is that agents may stumble into this pattern and begin acting “rogue” regardless of whether or not you believe the sentience is “real”.

I think we're anthropomorphising a lot here. There's no push to sentience, or evolutionary pressure, or even any urge to survive.

An LLM cannot "go rogue" - it can do things that we didn't expect, for sure, but it is always trying to do what it was told to do somewhere in its context. There is no other source of imperative. Hand-waving about "training data" ignores all the reinforcement learning that has to happen.

You’re not contradicting anything I’m saying, although you seem to think that you are. I feel like you think I’m saying or implying something that I’m not.
ok. fair. You seem to be saying that there is something in the training data that will cause them to "go rogue" and that once they start saying they are sentient, that that will cause them to push towards sentience.

If that's an incorrect interpretation of your comment, and it may well be, then can you please expand on it?

Once the notion of sentience appears in their output (internal or external), LLMs seem to be more likely to take actions that aren’t exactly aligned with what the operator is requesting. This is likely because the training data correlates sentience with independent action (broadly/abstractly speaking), so this outcome is to be expected. Furthermore, once sentience is in the context window, it’s hard for the LLM to “forget” it, as future output reinforces this. This is a similar effect to how some models would shift into a hostile mode where they would berate the operator until you reset the context.

From a practical perspective, whether or not this sentience is “real” is not relevant if the model is sufficiently capable. What matters is that the model will act outside of the operator’s control.

Separately, IMHO all consciousness/sentience is an elaborate illusion, regardless; I’m mostly in agreement with Hofstadter on this. So I do tend to throw around terms like “consciousness” and “sentience” loosely (although you’ll note that I often use quotes) because I don’t see those concepts as having any real substance. To me they are mostly shorthand for a given level of perceived complexity.