Hacker News new | ask | show | jobs
by ameliaquining 1 day ago
I'm curious, what are you hoping to convey by reminding people that LLMs are next-token predictors? They are, of course, but most people without an AI background won't fully understand what that means, so I assume you're using it at least partly as a proxy for something else.
1 comments

I think understanding how this stuff works is really important. For technical people it gives them a useful starting point for understanding it all. For less technical people it's crucial to help them understand that it's not some weird new magical science-fiction AI - it's still computer programs that turn text into numbers and do stuff with the numbers and turn those back into text.

It's harder to believe something is conscious or threatening to achieve word domination once you understand that it's a machine that statistically figures out which word should come next.

I don't think the next-token-predictor thing should increase anyone's confidence that LLMs aren't conscious or can't escape the control of their operators. A very closely analogous argument would "prove" that humans aren't conscious or can't do [insert task here] either. (No, I'm not saying that any of this is true of today's LLMs, I'm saying this particular argument doesn't work.)

I recommend this explanation: https://www.astralcodexten.com/p/next-token-predictor-is-an-...

You can say that for any argument regarding consciousness, because we don’t have an actual, all encompassing definition of what consciousness is. In general I don’t think comparison with humans makes much sense, we should be able to discuss LLMs without always falling back to “but what about humans” (sorry for the caricature)
Shouldn't that imply that agnosticism is the proper view, rather than asserting that something is impossible on a next-token-predictor architecture?

(Note: I don't actually think the consciousness question is the most important one in the near term. Where I think this line of reasoning gets really dangerous is when people use it to assert that LLMs can't or won't engage in certain behaviors no matter much they advance; this doesn't have anything to do with consciousness.)

If you believe that matrix multiplication with random sampling is conscious, then you probably believe everything is conscious, like rocks.

Most people would expect that matrix multiplication is not conscious, and autocomplete is not conscious either.

We can't prove matrix multiplication isn't conscious, but it doesn't seem likely unless everything is conscious.

There are some vague notions about consciousness emerging from complexity that some people may advance as a nuance to your point, although that just makes rocks “minimally conscious”, not necessarily unconscious. If you take the view that the entirety of our consciousness’s comes from purely classical interactions (electrical and chemical), then thats not a difficult conclusion to arrive at, but the notion that consciousness can be derived from deterministic computation does not pass the sniff test imo.

As a side note, that is why I find the idea that the brain is a quantum-classical hybrid computer appealing. And following the research developments is very interesting, to say the least.

Seems like a forest / trees error.

“If LLMs are conscious, it means matrix multiplication is conscious” == “If humans are conscious it means cells are conscious”.

It’s possible for complex systems to have emergent properties not exhibited by any individual component of the system.

I think you can reliably assert that X != Y without having a complete definition of Y, as long as you can identify at least one property or condition that Y possesses which X violates.

So for consciousness and LLMs it could be Qualia, lack of semantic understanding, lack of continuity in time, lack of a high degree of integrated causal feedback, etc.

Or perhaps those are just features of human consciousness but not integral to consciousness as a whole. To me this then implies panpsychism to some degree, which I'm alright with too.

I would think that qualia is the element of that list that actually matters here- "Is there a way that it feels to be an LLM?" is, I expect, the underlying question being asked within "Is an LLM conscious?". If LLMs could only experience the timeless, nonlocalized color blue, they'd still be conscious. The presence of positive and negative valence qualia is additionally relevant if someone is getting at whether they have moral worth, but it's secondary.

But qualia are not directly measurable and the rest of the list only matters if those features are necessary for qualia, which we can't decide without such measurements or at least a strong theoretical model.

I think it's important to understand the humans _can_ do what LLMs do: predict next tokens from prior ones.

But LLMs are only operating on text and humans are only operating on <waves hands>

While you’re active in this thread, I just want to say thank you for all your writing, you’re such a reliable source of sanity in that crazy new world :)
It's harder to believe something is conscious or threatening to achieve word domination once you understand that it's a machine that statistically figures out which word should come next.

The problem is, these models challenge our definition of "consciousness." Or at least they point out how hopelessly-inadequate our thinking on the subject is. Some people really, really don't like having their personal definition of consciousness challenged.

The correct response to "So what, it's just a next-token predictor" isn't a long dissertation on RLHF, training architectures, scaling laws and whatever, but rather to turn around and respond, "Sure, and how is that different from what we do?"

LLMs self-evidently have no experience when not inferencing, seems like the most consequential difference to me. If we are next token predictors then we are next token predictors that are inferencing at every waking moment and arguably much of our sleeping moments too; when we train an LLM that can inference as much as we do and remain coherent then I'll be more worried about whether it might have experience.
> It's harder to believe something is conscious or threatening to achieve word domination once you understand that it's a machine that statistically figures out which word should come next.

At the risk of sounding overly flippant, all world domination has been achieved by some person(s) figuring out which word should come next. Words quite literally = action when it comes to LLM’s with tools access

> At the risk of seeming X, statement that overwhelmingly demonstrates X-ness.

It's exhausting to even consider where to begin addressing the assertion that good leadership is just predicting the next word to say. Especially considering the corpus available to most great leaders in history was extremely small. To think Hannibal's military campaigns were just because he'd read like ten books in his life and could accurately forecast effective rhetoric is...indescribably divorced from reality.

At the risk of seeming like a jerk.

"Competent strategy" is just as good a target for a word predictor to optimize as "effective rhetoric" is.
I suggest you re-read what I wrote. In particular the part where I said there was no large corpus to perform this prediction back then.

After you edit, I'll address your point about how military strategy is nothing but skill with words.

> After you edit, I'll address your point

This has to be the most insufferable thing I’ve read/heard in weeks. Are you being serious right now?