| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by Trasmatta 409 days ago

I think we need to start moving away from this explanation, because the truth is more complex. Anthropic's own research showed that Claude does actually "plan ahead", beyond the next token.

https://www.anthropic.com/research/tracing-thoughts-language...

> Instead, we found that Claude plans ahead. Before starting the second line, it began "thinking" of potential on-topic words that would rhyme with "grab it". Then, with these plans in mind, it writes a line to end with the planned word.

3 comments

ceh123 409 days ago

I'm not sure if this really says the truth is more complex? It is still doing next-token prediction, but it's prediction method is sufficiently complicated in terms of conditional probabilities that it recognizes that if you need to rhyme, you need to get to some future state, which then impacts the probabilities of the intermediate states.

At least in my view it's still inherently a next-token predictor, just with really good conditional probability understandings.

dymk 409 days ago

Like the old saying goes, a sufficiently complex next token predictor is indistinguishable from your average software engineer

johnthewise 409 days ago

A perfect next token predictor is equivalent to god

lanstin 409 days ago

Not really - even my kids knew enough to interrupt my stream of words with running away or flinging the food from the fork.

Tadpole9181 408 days ago

That's entirely an implementation limitation from humans. There's no reason to believe a reasoning model could NOT be trained to stream multimodal input and perform a burst of reasoning on each step, interjecting when it feels appropriate.

We simply haven't.

lanstin 406 days ago

Not sure training on language data will teach how to experiment with the social system like being a toddler will, but maybe. Where does the glance of assertive independence as the spoon turns get in there? Will the robot try to make its eyes gleam mischeviously as is written so often.

jermaustin1 409 days ago

But then so are we? We are just predicting the next word we are saying, are we not? Even when you add thoughts behind it (sure some people think differently - be it without an inner monologue, or be it just in colors and sounds and shapes, etc), but that "reasoning" is still going into the act of coming up with the next word we are speaking/writing.

spookie 409 days ago

This type of response always irks me.

It shows that we, computer scientists, think of ourselves as experts on anything. Even though biological machines are well outside our expertise.

We should stop repeating things we don't understand.

BobaFloutist 409 days ago

We're not predicting the next word we're most likely to say, we're actively choosing the word that we believe most successfully conveys what we want to communicate. This relies on a theory of mind of those around us and an intentionality of speech that aren't even remotely the same as "guessing what we would say if only we said it"

ijidak 408 days ago

When you talk at full speed, are you really picking the next word?

I feel that we pick the next thought to convey. I don't feel like we actively think about the words we're going to use to get there.

Though we are capable of doing that when we stop to slowly explain an idea.

I feel that llms are the thought to text without the free-flowing thought.

As in, an llm won't just start talking, it doesn't have that always on conscious element.

But this is all philosophical, me trying to explain my own existence.

I've always marveled at how the brain picks the next word without me actively thinking about each word.

It just appears.

For example, there are times when a word I never use and couldn't even give you the explicit definition of pops into my head and it is the right word for that sentence, but I have no active understanding of that word. It's exactly as if my brain knows that the thought I'm trying to convey requires this word from some probability analysis.

It's why I feel we learn so much from reading.

We are learning the words that we will later re-utter and how they relate to each other.

I also agree with most who feel there's still something missing for llms, like the character from wizard of Oz that is talking while saying if he only had a brain...

There is some of that going on with llms.

But it feels like a major piece of what makes our minds work.

Or, at least what makes communication from mind-to-mind work.

It's like computers can now share thoughts with humans though still lacking some form of thought themselves.

But the set of puzzle pieces missing from full-blown human intelligence seems to be a lot smaller today.

thomastjeffery 409 days ago

We are really only what we understand ourselves to be? We must have a pretty great understanding of that thing we can't explain then.

mensetmanusman 408 days ago

I wouldn’t trust a next word guesser to make any claim like you attempt, ergo we aren’t, and the moment we think we are, we aren’t.

hadlock 409 days ago

Humans and LLMs are built differently, it seems disingenuous to think we both use the same methods to arrive at the same general conclusion. I can inherently understand some proofs of pythagorean's theorem but an LLM might apply different ones for various reasons. But the output/result is still the same. If a next token generator run in parallel can generate a performant relational database that doesn't directly imply I am also a next token generator.

skywhopper 408 days ago

Humans do far more than generate tokens.

Mahn 409 days ago

At this point you have to start entertaining the question of what is the difference between general intelligence and a "sufficiently complicated" next token prediction algorithm.

dontlikeyoueith 409 days ago

A sufficiently large lookup table in DB is mathematically indistinguishable from a sufficiently complicated next token prediction algorithm is mathematically indistinguishable from general intelligence.

All that means is that treating something as a black box doesn't tell you anything about what's inside the box.

int_19h 409 days ago

Why do we care, so long as the box can genuinely reason about things?

chipsrafferty 408 days ago

What if the box has spiders in it

dontlikeyoueith 408 days ago

:facepalm:

I ... did you respond to the wrong comment?

Or do you actually think the DB table can genuinely reason about things?

int_19h 408 days ago

Of course it can. Reasoning is algorithmic in nature, and algorithms can be encoded as sufficiently large state transition tables. I don't buy into Searle's "it can't reason because of course it can't" nonsense.

Tadpole9181 409 days ago

But then this classifier is entirely useless because that's all humans are too? I have no reason to believe you are anything but a stochastic parrot.

Are we just now rediscovering hundred year-old philosophy in CS?

BalinKing 409 days ago

There's a massive difference between "I have no reason to believe you are anything but a stochastic parrot" and "you are a stochastic parrot".

ToValueFunfetti 409 days ago

If we're at the point where planning what I'm going to write, reasoning it out in language, or preparing a draft and editing it is insufficient to make me not a stochastic parrot, I think it's important to specify what massive differences could exist between appearing like one and being one. I don't see a distinction between this process and how I write everything, other than "I do it better"- I guess I can technically use visual reasoning, but mine is underdeveloped and goes unused. Is it just a dichotomy of stochastic parrot vs. conscious entity?

Tadpole9181 408 days ago

Then I'll just say you are a stochastic parrot. Again, solipsism is not a new premise. The philosophical zombie argument has been around over 50 years now.

dontlikeyoueith 409 days ago

> Anthropic's own research showed that Claude does actually "plan ahead", beyond the next token.

For a very vacuous sense of "plan ahead", sure.

By that logic, a basic Markov-chain with beam search plans ahead too.

cmiles74 409 days ago

It reads to me like they compare the output of different prompts and somehow reach the conclusion that Claude is generating more than one token and "planning" ahead. They leave out how this works.

My guess is that they have Claude generate a set of candidate outputs and the Claude chooses the "best" candidate and returns that. I agree this improves the usefulness of the output but I don't think this is a fundamentally different thing from "guessing the next token".

UPDATE: I read the paper and I was being overly generous. It's still just guessing the next token as it always has. This "multi-hop reasoning" is really just another way of talking about the relationships between tokens.

Trasmatta 409 days ago

That's not the methodology they used. They're actually inspecting Claude's internal state and suppression certain concepts, or replacing them with others. The paper goes into more detail. The "planning" happens further in advance than "the next token".

cmiles74 409 days ago

Okay, I read the paper. I see what they are saying but I strongly disagree that the model is "thinking". They have highlighted that relationships between words is complicated, which we already knew. They also point out that some words are related to other words which are related to other words which, again, we already knew. Lastly they used their model (not Claude) to change the weights associated with some words, thus changing the output to meet their predictions, which I agree is very interesting.

Interpreting the relationship between words as "multi-hop reasoning" is more about changing the words we use to talk about things and less about fundamental changes in the way LLMs work. It's still doing the same thing it did two years ago (although much faster and better). It's guessing the next token.

Trasmatta 409 days ago

I said "planning ahead", not "thinking". It's clearly doing more than only predicting the very next token.

therealpygon 409 days ago

They have written multiple papers on the subject, so there isn’t much need for you to guess incorrectly what they did.