Hacker News new | ask | show | jobs
by Retr0id 7 days ago
Until recently I thought "load-bearing seam" was a satirical exaggeration - I'd seen both claudisms independently but never combined. But a couple of days ago it hit me with "The key structural point first: the only load-bearing seam is [...]"
7 comments

It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.

It was rewarded for this during training for some reason.

Alternative theory:

The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans.

But humans generally prefer to communicate on low levels of abstraction, through a back and forth, until the hard-to-express higher abstraction exists in the head of everyone involved - without ever being directly communicated. This is because we don't think using language. Language is merely a lossy translation of our thought into something expressible, happening after the fact or alongside it.

So when the LLM starts speaking to us using patterns and terms it created for itself during training to encode abstract thought in language, communicating with it becomes painful.

"There is a depth of thought untouched by words, and deeper still a depth of formless feeling untouched by thought." - Rilke

Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language forms an essential part of thinking beyond a base layer of instinctive animalistic associations. Sophisticated thoughts are impossible to construct or maintain in the absence of language to represent concepts.

Edt to add: I cited Rilke because I find the notion [some deep thoughts are beyond language] interesting. But I disagree with the idea that language is only ever epiphenomenal (co-occurring with thought), or akin to a hard-of-hearing scribe attempting to convey thoughts which always have independent existence.

I strongly disagree. For one thing, many animals that lack language can still navigate a very complicated natural world using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things. Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.

This goes back to the whole Tarzan obsession of the early 20th, I guess, or earlier. But we know that apes can make simple tools without the language to describe them, or the thought process that went into them.

Thought is multimodal. Language is just one lossy mode.

> For one thing, many animals that lack language can still navigate a very complicated natural world

So can a cruise missile. Also I think there's like separate part of the brain for that

> using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things.*

FWIW, AFAIK we haven't shown the ability to think in concept exists anywhere except in humans (because philosophy, reported experience) and LLMs (because we can literally see them forming and activating in patterns, and we've learned to identify them specifically, and experimentally verified through amplifying or suppressing them and observing behavior, etc.).

But more importantly:

> Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.

Images and music and math are langauge. If it can convey ideas, it is language.

Words and sentences and speech are subset of the idea of language and communication, that for some reason gets routinely confused for the whole thing. At this point I'd say even the "language models" are badly named, simply because people see "language models" think of "token" as number representing a sub-word element in existing human language like English. With multimodal models, at this point tokens are closer to units of sensory experience.

Have you ever had a professor or mentor walk you through some experiences as a way to help you understand some complex concepts? I guess language is used in that but the ideas are conveyed by the process not the language.

And I rather disagree that mathematics is language, when I did maths there was a distinct difference from understanding a thing and then writing it down.

And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.

I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen. In the meantime, in my hands, they can do very nice dataviz programming to make complex systems much more transparent, even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.

> And I rather disagree that mathematics is language, when I did maths there was a distinct difference from understanding a thing and then writing it down.

It is generally the case that there's a difference between "understanding a thing" and "writing it down", as demonstrated by every student who studies for the exam, the Chinese Room thought experiment, and Business Bullshit Bingo.

For example, I can copy the next sentence of yours, but I have no idea what:

> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.

means, can you rephrase that by as much as possible? Preferably without a double negative?

Agree regarding sanity issues of AI. Extremely unlikely we make something "stable" so early in our attempts.

> Have you ever had a professor or mentor walk you through some experiences as a way to help you understand some complex concepts? I guess language is used in that but the ideas are conveyed by the process not the language.

Yes, but I believe this is happening with LLMs too. Language (as in text, symbols, actions) is a vehicle, but much like us, language models have internal models and represent concepts (this has been directly, empirically demonstrated few years ago), and they don't "think" in tokens either[0].

> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.

That's fair. LLMs don't capture every dimension we experience. There's also history - our individual lived experiences since birth are, in my view, something else entirely. It's not a modality, but it's also not something currently possible to capture in or post training.

> I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen.

Possibly. I definitely agree it's not human intelligence. I think it's human-like, in the sense of human-approximating, by virtue of how it's trained[1], but it's arriving there via a different path so end result can still be alien (though I speculate approximation will hold[2]).

> even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.

I guess depends on the complexity of the case (and my understanding of your example); e.g. in my case, I had stellar results from giving Sonnet and Opus (from 4.6 all the way to now) access to my Home Assistant instance. They aren't perfect at making dashboards, but they're excellent at surfacing insights I didn't even realize were possible to get.

--

[0] - That's distinct from the "tokens are units of thinking" heuristic, which still holds for mechanistic reasons - best analogy IMO is "clock signal" in ICs.

[1] - The overall goal function is literally just "generate output that looks sensible to a human", in fully general, unqualified sense. Or, put another way, we're just brute-forcing DWIM, and rating the output by whether it's "what I meant".

[2] - Thinking about constraints on biological evolution, whatever the design of a human mind is, the fundamentals behind it must be so simple, that a greedy incremental optimizer random-walked into it. Given how far we've got with LLMs using simple architecture and crude training methods, and especially how eerily similar their failure modes are to human cognitive failure modes, I suspect LLMs are actually attracted towards the same fundamental design as evolution discovered.

Ok, but it is clear the parent meant language as in words like english or tamil. If you expand language to include anything that can convey idea then the argument is meaningless and its a null agreement.
It's not a null argument, because LLMs have long moved past "language as in words like english or tamil". Tokens are not units of such language, at least not since visual and audio tokens became first-class objects and there's no actual distinction between them inside the model. That sameness of representation is what it means for a model to be "multimodal".
This is true and it's barely even debatable. Whatever exact role language plays in our thought processes, it is most definitely nonzero.

It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks.

Linguistic Relativity — John Lucy https://www.annualreviews.org/doi/10.1146/annurev.anthro.26....

Russian Blues Reveal Effects of Language on Color Discrimination https://www.pnas.org/doi/10.1073/pnas.0701644104

Unconscious Effects of Language-Specific Terminology on Pre-Attentive Color Perception https://www.pnas.org/doi/10.1073/pnas.0811155106

Newly Trained Lexical Categories Produce Lateralized Categorical Perception of Color https://www.pnas.org/doi/10.1073/pnas.1005669107

> sort of crystallized a bit of the human thought process

a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there.

b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.

When I see these sorts of debates about LLMs thinking, its rarely a disagreement about what LLMs do. Its almost always over how 'thinking' is defined and the two sides use different definitions but don't actually communicate to each other what those definitions are because they assume the other side is using the same one.

The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.

It has nothing to do with thinking or consciousness.

There is a common misconception that LLM are simply a "statistical process" that doesn't feature any abstract conception of the tokens it is predicting. There are studies that show that such features do exist - that there is discernible structure built into the weights - and that the process of inference is a very rich one.

The statistical process exists but it is the substrate in which the model is implemented - or more accurately - grown.

If you can predict Magnus Carlsen's next move then you are just as good at chess as Magnus - and being that good absolutely does require reasoning.

If you can predict the solution to an open Erdos problem that stumped hundreds of people for decades...

   When I see these sorts of debates about LLMs thinking, 
   its rarely a disagreement about what LLMs do. Its almost 
   always over how 'thinking' is defined and the two sides 
   use different definitions
Well, hmmm. Yes, I think that happens a lot.

I think there's a pattern that happens even more often, and it's what happened here.

Whether I'm right or not, what I said was somewhat nuanced - I stated language is a part of our thought process (even posted research to support this) and, given that fact, I think many underrate how wild this achievement is even if it's only "fancy autocorrect."

And, of course, the other side comes in with BUT IT'S NOT THINKING.

Which... I didn't say, and I would not say, because (like you said) it's impossible to do without the discussion immediately devolving into semantics. Semantics that I'm really, really uninterested in. But, FWIW, I like your definition.

Did you reply to the right post? You quoted me, but you wrote "LLMs don't think" as if it was a rebuttal. It's puzzling, because I didn't say that they think, so it kinda seems like you got confused? Maybe somebody else said that?

I don't really have an opinion on whether or not they "think" because I feel it's impossible to even discuss without getting into a very very uninteresting semantic argument about what "thinking" is.

Are we defining "thinking" as doing it the same way humans do it? Then, of course they're not thinking. It's a statistical model, not axons and neurons, or even a simulation of axons and neurons.

Are we defining "thinking" on a purely functional or behavioral basis, kind of a Turing test approach? Then... well, I think it gets nuanced. For some tasks, within some constraints, they do pass that test. For many others, of course they don't.

Are we defining thinking in more esoteric terms? Something to do with the soul? Maybe the ability to come up with truly novel concepts rather than rehashing and remixing the stuff it was trained on? Do ants think? Do dogs think? Do jellyfish think? Octopi? A newborn baby?

Anyway, it's a deeply uninteresting semantic question.

You should look into emergence. An ant in a colony, an offer in a market, a drop of water in a weather system are all evidence that irreducibly complex things have simple mechanisms at their core.
The (or a) current neuroscience models of the brain are that its main job is to predict how the body should be responding in the near future. Obviously a lot more complex network nodes than an LLM, but prediction is clearly tied up with thought in some way.

I don’t find LLMs to be very good independent thinkers, but I wouldn’t over sell our own mentation either - it clearly arises from a large number of simpler entities.

The more significant difference is that the LLM is stuck with language which is clearly an emergent and secondary capability of our own thinking. We can formulate words to explain things, but we also can look at two volumes and feel what it means that one is larger than the other. Raise a toddler and you can see the progression from not understanding, repeated experiments, muscle memory and finally to conscious point for reasoning.

    I don’t find LLMs to be very good independent 
    thinkers, but I wouldn’t over sell our own mentation 
    either - it clearly arises from a large number of 
    simpler entities.
Yeah. I don't see them ever hitting the heights of human creativity in terms of coming up with entirely new ideas, schools of thought, etc. That really might be a fundamental limitation of being trained on existing thought. Also, a lot of human experience involves (1) things we don't have words for (2) things we've never put into words.

    clearly an emergent and secondary capability of 
    our own thinking.
Yes. And it's part of our thinking. More than a capability . Thought influences speech, but speech also influences thought.

That's why I think it's remarkable that we've managed to (choosing my words very, very carefully here) create a statistical model that does a remarkably decent job at emulating the behavior of a fragment of that process.

The statistical model that underpins any deep learning system is the substrate in which a process is implemented. There is still a process - just because that process is grown - not programmed - doesn't tell us anything about the depth or limits of its capability.
It's amazing that you can predict a counterexample to an open math problem, all without thinking.
Yet they do.
Not really, a shitload of "open math problems" are bounded by constraints of simple text processing or heuristic search.

Much of math is just boring routine work.

"Fancy logistic regressors" are, in fact, a modelling tool. You can tell LLMs model human behaviour because they're doing things that until a few years ago only humans could do, like cheat in exams.
When I was young, a friend asked me, "Hey, you speak three languages, which one do you think in?"

I paused, confused, and replied, "People think in words?"

Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)

In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.

More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.

The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.

Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)

I find being alone, with no other human around for miles, gets rid of language. Same with combat sports, and more broadly, struggling on the edge of my physical ability.
You just lack (or lacked) introspection, nothing special. Many people claim to have no inner monologue. When pressed, it always comes out. Simply put, it's impossible to function as a human without it. But many people are not aware, and think it means something like "hearing actual voices".
"Thought is language" is a folk intuition that has been proven wrong over and over. People with severe global or agrammatic aphasia still can score higher on intelligence tasks than you'd expect - even though this data generally comes from people who suffered a stroke, which generally isn't neatly contained to one's language capability. Also, more importantly, even in people with full brain function, the language network appears remarkably unresponsive under neuroimaging when performing arithmetic, logic puzzles, or theory-of-mind tasks. Also famously music tasks: Vissarion Shebalin lost most of his language capabilities to a stroke and still completed his fifth symphony.

There's correlation, and your language network likely augments your intelligence, but as proven by millions of animals, unfortunate humans and also some less-unfortunate human infants, you really don't need language for intelligence.

Where language helps most strongly is metacognition: evidence suggests that it is severely limited without language, and I suppose that is where the common belief that thought is language comes from: the moment you try to think about your thoughts, you use language!

My theory (and this with literally no evidence) is that we use the language network for metacognition precisely because it is not that involved in primary thought. Important to note that "we" here means humans: animals appear to demonstrate metacognition even without language.

Isn't the folk intuition that "words come after thoughts"?

I agree that a stroke cannot isolate language exactly, it's not an ideal approximation of "languagelessness". I think you will find this report interesting: https://nautil.us/what-my-stroke-taught-me-236544

There is no doubt that many marvelous works and concept can be produced entirely without language. Something as complex as driving a car can be done mostly without "thinking about it", without language, without the inner monologue going over all the decision our mind is making.

What we can't do without language, is making a complex plan over time and space, beyond something like "need to fetch hidden item behind the corner to unlock box right in front of me now". The question is how much exactly does language "augment" our intelligence. Is a baby, a human that has not yet learned language, really smarter than an octopus? Or a crow, remarkably smart but interestingly not a mammal, which indicated that intelligence doesn't have to be strongly tied to some evolutionary biological feature. A human stays as "dumb" as an animal if they don't learn language.

Re: your theory, if true, why? The theory implies (though correct me if I'm wrong and strawmanning), that there is another, hidden, way of encoding these extremely complex "information bundles" for lack of another word, in our brains. Why not use that for metacognition too, then?

When I learned French as an adult, I pushed myself heavily into "forced immersion", avoiding English completely for long periods of time as my French slowly built itself up. This almost entirely silenced my inner monologue in English for those periods of time, and left me with a toddler's ability to think in French. In this comparative silence, it was much easier to observe my non-verbal thought processes moving around "behind" the scarce words.

Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously translate from French to English for about 15 minutes until my brain switched back.

So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.

Also, something which at least some other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.

You just stated your vibes like it is a fact.

There was a study about this: https://journals.sagepub.com/doi/abs/10.1177/095679762412430...

> adults who reported low levels of inner speech
I'm not sure. I think I had a reasonable amount of metacognition as a child. I remember being eight years old, and having a really cool idea but forgetting it.

So I devised a plan to retrieve it. I'm just going to rewind time. I'm going to go back to doing what I was doing when I had that idea, and then it'll come back to me.

And so I remembered that I had been fiddling with my seat belt when I had the idea. And so I resumed fiddling, and my cool idea promptly came back.

Metaphors and abstractions came to me much later though. (I struggled with OOP for about 10 years until one day it all just clicked.) Symbolism took me another ten years.

> So I devised a plan to retrieve it. I'm just going to rewind time. I'm going to go back to doing what I was doing when I had that idea, and then it'll come back to me.

You narrated the ideas back. It's words and language. There is no hidden, unknown, layer of "thought".

Language is a reductive and lossy serialization of "thought-stuff". Sometimes you need this information-shedding to clear your working memory to make room for more things. Sometimes it's literally just a way to communicate. You're turning something fuzzy into something discrete.

E.g. "I'm feeling something. Is it anger? Yes, I'm angry." But in reality anger isn't just one thing. It's a cluster of infinite and varied feelings that we label as "anger". Something is lost when we do this labeling.

Notice then that the feeling of "anger" didn't start from your language, you merely used language to label, discretize, classify, standardize, compress it. It's one-way.

Also there are emotions for which society has not developed words because it has a social tendency to deny them. OCD, ADHD and Tourettes, for example are driven by deep, complex, strange destructive emotional needs that are wordless and undescribed. They are reduced to “obsession”, “compulsion”, “urge”, “tic”, etc., words that essentially only describe the appearance of the outcomes to others, which are a completely hollow description of the extraordinary internal experiences.

The very fact that we don’t have words for such powerful internal experiences is one of many reasons I find the LLM enthusiasts’ belief that LLMs will one day write indistinguishably from humans to be hollow.

To me, Alexithymia alone proves this.
There are people who do not have any words in their heads at all when they think, and there doesn't seem to be any reason to disbelieve them. It may have been proven or at least observed in fMRI. Some people think in visuals.

I think some of this was discovered somewhat recently

I think you're conflating different things. Many people do not have an internal 'narrator'. I am one. I don't have a voice in my head saying words, ever. I do definitely have something like a playback of other people saying things, though. Words are still in there in the form of recall, they are just not part of the executive layer in a way I have access to.
What about when you're reading? You don't have a voice in your head saying the words?
Only if it is a very hard text that I have to move slowly thru. For normal reading, no not at all. If I am trying to read something faster than is comfortable, I will periodically emphasize the thing, the idea, behind a key word, so I am more likely to remember it at the useful time. Not verbal and there are many many times when I can describe the thing but not recall the normal word. Like the flat thing with keys in it instead of keyboard. Or even weird quasi-synonyms that serve to muddy the waters as far as sharing my thoughts with others goes.
No. Do you? Like you hear words in your head?

Obviously there is a language layer in there somewhere, but my 'driver' doesn't access it. I don't think in words, I don't have an internal narrative, and when I'm reading I don't even 'see' individual words, in the same way you aren't thinking about your ankle muscles when you're hitting the brakes while driving.

I definitely can consciously make words, but it's 'me' deliberately deciding to make them and hold them in my head. It wouldn't happen naturally.

And some people can’t see any pictures in their minds! My partner is one of these people. We both really enjoy reading fiction books, often with a fantasy or sci-fi bent and it genuinely amazes me that they can experience these books in a way that feels wholly alien to me.

The fun of reading to me is constructing the world in my minds eye and turning the words on the page into a visual experience only found in my mind using imagination. This is a reason why many people get upset when a movie adaptation is made and the actor chosen for their favorite character feels very off or wrong; their mental picture of that character is totally different and it causes dissonance that our brains don’t like. For my partner this is a non issue because they never make a mental image of the person, so the movie is genuinely the first time they are “seeing” a physical representation of the character.

The human mind is genuinely amazing and fascinating and I believe that this range of human experience will be the final 20% for “AI” that might never be reproducible.

Does anyone else just end up with a random celebrity 'playing' characters in your brain image of scenes? I find that I generally picture a person that I have seen on TV, with details that don't nessesarily match book character details.

For example, I recently read a novel where there is 'tough lady cop' character and in my brain, entirely involuntarily, the role has been assigned to Brooklyn Nine-Nine's Rosa Diaz in exactly the clothing and context she exists in that TV show.

When the character in the book has a certain clothing or whatever it all kinda glosses over until it's back to the character I have seen before.

I don't know if this is more a male thing or am ADHD thing or what but details of characters dress and looks are pretty much lost on me once there are assigned a character from 'central casting'

> This is a reason why many people get upset when a movie adaptation is made

I still haven't watched the Dune movies because of this reason. I liked the book a lot, and had a very personal image of what the world looked like, and was afraid to lose it. Unfortunately, by now, I've seen many video clips on social media, and my internal imagery has already been poisoned. Might as well watch the movies at this point.

This is highly contested, most likely not true, and probably not measurable anyway.

https://journals.sagepub.com/doi/10.1177/09567976251335583

Just because you can think through writing doesn't mean language is the essence of thought itself.
The virtue of writing is it makes it harder to fool yourself that you have all the important links addressed in your construction/argument/proposal.
> Your assertion that we don't think in language is questionable.

It's true though; we routinely see people get stuck for a word that they know but can't quite recall at that moment in time. It happens daily across billions of people, yourself included.

If we thought in language, it is impossible to be stuck for a specific word. But we all experience this at some point in our lives, hence we aren't thinking in language.

> It runs counter to the lived experience of developing thoughts through writing

Writing requires thought, but writing isn't thought. Just like doing requires thought, but doing isn't thought, writing is a subset of doing.

You can also doodle to think things through, or play with toys to think things through, or many other similar things.

Hand-waving assumptions drawn from narrow subjective experience without awareness of the profound and contradictory results that have come from scientific study of consciousness and causality.
I'm not an expert but I saw Yann LeCun shared this recent Neuroscience article on his Facebook page commenting "I don't think in words. Animals don't think in words."

Evidence from formal logical reasoning reveals that the language of thought is not natural language

https://www.pnas.org/doi/10.1073/pnas.2520095123

Are you sure this is Rilke? Couldn’t find anything online - it seems it’s close to a quote by Zora Neale Hurston
The term 'language model' throws some people off thinking you can only put english or french, or both into a model. Technically an LLM can learn about anything that can be digitized. If you wanted to spend a billion dollars training one on wireless signals it wouldn't be impossible for it to connect to your router with the right antenna attached. So only limiting it to the idea of language leaves off a lot of other types of abstractions and concepts they encode.
shhh don't give anyone any ideas, because they'll try it and sell it to management, as idiotic as the idea itself is
I have a fantasy that we will do this to turn trees into antennas again, with usable bandwidth.
> even when the claims aren't based on anything and are possibly wrong.

Are you saying Claude is engaging in Rhetorics because the RL data generated by humans were influenced more by it and persuasion rather than actual logic or reasoning?

I think that's precisely what they're saying. It shouldn't be a surprise that it's successful. Eliza proved the same thing 35 years ago.
It's actually an effect that happens in the (re-)alignment process due to harmonic properties of the positional encoding in the attention matrix.

(I recommend reading and implementing the Attention is all you need paper. By hand. Otherwise you won't learn anything from it.)

> It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.

When your training set contains more or less the complete output of every capital-C Consulting firm...

Maybe we need LLMs which have an internal dialogue rather than the current monologue.
> Maybe we need LLMs which have an internal dialogue rather than the current monologue.

We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.

All of these rhetorical devices to make the assertion seem authoritative and correct are derived from academia
You can thank everyone who thumbs up this response along the training pipeline.

Modern LLMs are not too different from Reddit/Twitter in that regard, I'm sure the AI labs learned (lol) a lot from them re: how to do "engagement".

This matters. That's the spine of it.
the phrase load-bearing caveat makes me irate
Only that one? Not belt-and-suspenders enough
It's justa caveat is side to the main point. No human would say that. It's like a load bearing utility shed
People make fun of the language, rightfully so in some cases, but also it's often quite effective language. "load-bearing seam" communicates quite a lot in very few characters.
It's a clunky metaphor, and it's used clumsily: not for the sake of its clunkiness, but by mixing two decent individual parts in an attempt to have them reinforce each other, but ending up with something weaker.

It communicates an absence of thought and awareness, blind groping at building blocks without understanding. It's borderline vapid, and quite annoying.

I don't follow, the words make perfect sense together in most software engineering contexts.

"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.

To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.

A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.

Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.

In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.

Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.

Well, first, we don't have enough context to judge what the "load-bearing seam" was used for in this case ("The key structural point first: the only load-bearing seam is [...]"), so we don't know if it was meant in the well-defined legacy-code sense. But that doesn't really matter.

I'm taking issue with the combination "load-bearing seam". It's a bad metaphor, because seams are usually structural weak points in the physical world, and not load-bearing in the sense that this modifier is usually used. (Seams need to bear loads and stresses to do their job, but so do walls; yet, we do not call all walls load-bearing. We mean something extra when we say that, something that seams don't do.) Even if we were talking about seams in the well-defined software sense, as opposed to the metaphorical one, you still get a mixed metaphor as a result that I find extremely awkward and grating. It doesn't have to be. There are so many ways to highlight the importance of something without calling it "load-bearing".

I understand that you don't see it that way or don't care, but to me, the result is thoughtless, careless and vapid. Bad metaphors put little holes into a text, they leave eddies of confusion where meaning should be, they look load-bearing while actually being weakening, they're like a fart in the elevator that should lift the reader's understanding.

Note that I'm not calling into question that seams may be well-defined in some software contexts, or that "seam" and "load-bearing" can be valid metaphors on their own, as you describe. I think you might have misunderstood me that way. I'm only calling out, and fed up with, the bad style that permeates LLM-generated prose like the whiff of something not quite digested.

It's not this particular case that irks me, but what it exemplifies. I wouldn't mind so much if similar things to this weren't there everywhere, every single day.

If it doesn't bother you, I'm happy for you.

Or in a way, nothing at all.
On its own, yes, but not in context. I think one of the problems with Claude's stock output is it assumes the reader has a firm grip on the context of the output, which is often untrue. Stock output is exhausting to read because you have to unwind metaphors in an unstated context.
What I’m saying is that it often uses metaphors in ways that are subtly wrong or don’t make any sense.
I see it at least 3-4 times per week...
Fable has a Chomsky grammar.
Hello, completely hijacking your recent comment because I found no other way to contact you.

You wrote a month ago that you were close to releasing something called "nix-compile", how is it going ?