|
|
|
|
|
by delis-thumbs-7e
22 days ago
|
|
I am no linguist, but I believe this is referred to as surface structure and deep structure. What you describe is a line of text that is grammatically somewhat adequate to pass as readable and you treat all text the same. when we read the text, we decipher meaning out of according to word references and syntax - just like with Python or C++. However, if this were solely the case, we could not read Finnegan’s Wake. Probably vast majority of modern poetry would be unread by anyone, as would be pretty much all major works of philosophy Kant onwards. Deep structure is what according to Chomsky et al. gives meaning to the language, ie. somewhat logical structure behind the mere words. English word strict has order, need it but actually not does. You skibidi rizz swag grok brah also, barely. We use language in the extended meaning of the word to create a model of the world and somehow the past riverrun skibidi transmits that model to others. This is what I believe Bender also tries to say in their paper. Now, you could claim that LLM’s have this deep structure, create models of the world and are basically just like us, and certainly many here are adament that this is the case, being aghast how someone can “insult” LLM’s by calling them parrots. However, there really is not much proof to back up this belief. Usually LLM’s seem to copy existing surface structure from whatever source, and when it deviates from these patterns, it usually becomes incomprehensible. There’s much hoopla about LLM’s solving hard maths, but it seems even there they are mostly generalising from vast amounts of training data, rather than actually reasoning: https://arxiv.org/pdf/2410.05229 |
|
But even here, in the Chomsky sense, LLMs clearly exhibit deep structure because they can write at length in an internally consistent manner. Importantly, early generations, GPT-2 and even GPT-3, did not definitively have this property; roughly, an object that was green at the beginning of a paragraph might not still be green at the end of the paragraph. This was strong evidence for lack of a world model.
Current LLMs do not show this behavior. We cannot prove that LLMs have a world model, in fact, their architecture seems to rule it out, but looking at it from a linguistic standpoint, they produce language in a manner as if to reflect a world view. That is, we cannot easily falsify the statement "LLMs somehow represent a world model"; and current examples of "disproving" their world view are so convoluted that even humans do not appear to (observationally) have a world view either.
I'm not making a claim as to LLMs having genuine deep structure or consciousness or anything like that. I'm claiming that we can't rule out current or future capabilities or make structural assumptions. Yes, they generalize from their training data, but unless you can make very specific claims about the kinds of things that they _cannot_ do, I can't take this statement as particularly compelling.