Hacker News new | ask | show | jobs
by the8472 6 days ago
Yeah that's pretty much what gwern argues here[0]. Or to adapt another proverb: to predict the next token you first need to model the universe.

[0] https://gwern.net/scaling-hypothesis#gwern-difference--effic...

1 comments

> to predict the next token you first need to model the universe

Exactly. The "most likely next" series of tokens, for example, when given the first half of a correct mathematical proof, is the correct rest of the proof. I have never seen anyone define "most likely next token" in such a way that this isn't true.

It's like saying that thinking cannot generate new knowledge because all thought is just rearranging the information we get from our senses, or memory of previous information from our senses, and we do nothing more than figure out the most likely word to say next in a conversation.

Either humans are not capable of intelligence or computers are capable of becoming intelligent. Neither or both.

i think people say that thinking that only training to produce the next likely word would end up producing some local minimum word that generally fits but doesn't actually lead to intelligent thought.

that feels like a misunderstanding of how the loss function behaves when used within a sequence