Hacker News new | ask | show | jobs
by voidhorse 13 days ago
I'm really starting to tire of people making broad, general claims about how LLMs work or how to use them with N = 1 or 2.

An LLM is a statistics machine for goodness sake. Basically any general claim about them needs to exploit the law of large numbers to be even remotely sensible. You cannot extrapolate from one-off behavioral successes. LLMs are not understanding anything in the way humans do. If they did, yeah, maybe you could extrapolate hard from small samples, but they don't work or understand things like we do. You need to show that the behavior you are documenting is an average behavior the LLM converges toward in the long run.

1 comments

you have a premise at the heart of that:

> understanding anything in the way humans do

i'm not sure it's clearly established LLMs can't be a model of some part of “the way humans do”?

to be more specific, i'd argue LLMs “understand” awfully similarly to a brilliant (polymath) with early dementia or Alzheimer's

no executive function, no short term memory, and absent both of those, conversing with that person about the past or with an LLM about topics that had been "in their training sets before a cutoff date" is surprisingly similar, right down to: introing a topic precisely the same way, you'll experience the same conversation; convo loops if bits are too quantized (looping and lossiness / recall / context-length are correlated); and ofc opening a new session is like the first one never happened

> to be more specific, i'd argue LLMs “understand” awfully similarly to a brilliant (polymath) with early dementia or Alzheimer'

At which point i'd argue that you're kind of refuting your own claim as most humans are not brilliant polymaths with early onset alzheimer's.

Even admitting these kinds of comparisons to edge case human mental experience, there are more differences than similarities, and the similarities are misleading and superficial. There are still a lot of differences, even when it comes to the physical structure of the brains neurons as compared to digital neural nets, and i think it's far more beneficial to try to understand these machine in their uniqueness and for what they are than to draw hasty and shallow comparisons. You could be forgiven for that, though, because computing is rife with people who love to draw hasty unjustified analogies for some reason (example: people were already likening the brain to a computer when we didn't even have working implementations of neural nets yet and a computer was literally just a small number of logic gates lol)