Hacker News new | ask | show | jobs
by hmokiguess 7 days ago
Now I have the full picture. You're right to push back, and that's on me. The load-bearing seams of language are the smoking gun I should have been aware of.
9 comments

Until recently I thought "load-bearing seam" was a satirical exaggeration - I'd seen both claudisms independently but never combined. But a couple of days ago it hit me with "The key structural point first: the only load-bearing seam is [...]"
It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.

It was rewarded for this during training for some reason.

Alternative theory:

The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans.

But humans generally prefer to communicate on low levels of abstraction, through a back and forth, until the hard-to-express higher abstraction exists in the head of everyone involved - without ever being directly communicated. This is because we don't think using language. Language is merely a lossy translation of our thought into something expressible, happening after the fact or alongside it.

So when the LLM starts speaking to us using patterns and terms it created for itself during training to encode abstract thought in language, communicating with it becomes painful.

"There is a depth of thought untouched by words, and deeper still a depth of formless feeling untouched by thought." - Rilke

Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language forms an essential part of thinking beyond a base layer of instinctive animalistic associations. Sophisticated thoughts are impossible to construct or maintain in the absence of language to represent concepts.

Edt to add: I cited Rilke because I find the notion [some deep thoughts are beyond language] interesting. But I disagree with the idea that language is only ever epiphenomenal (co-occurring with thought), or akin to a hard-of-hearing scribe attempting to convey thoughts which always have independent existence.

I strongly disagree. For one thing, many animals that lack language can still navigate a very complicated natural world using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things. Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.

This goes back to the whole Tarzan obsession of the early 20th, I guess, or earlier. But we know that apes can make simple tools without the language to describe them, or the thought process that went into them.

Thought is multimodal. Language is just one lossy mode.

> For one thing, many animals that lack language can still navigate a very complicated natural world

So can a cruise missile. Also I think there's like separate part of the brain for that

> using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things.*

FWIW, AFAIK we haven't shown the ability to think in concept exists anywhere except in humans (because philosophy, reported experience) and LLMs (because we can literally see them forming and activating in patterns, and we've learned to identify them specifically, and experimentally verified through amplifying or suppressing them and observing behavior, etc.).

But more importantly:

> Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.

Images and music and math are langauge. If it can convey ideas, it is language.

Words and sentences and speech are subset of the idea of language and communication, that for some reason gets routinely confused for the whole thing. At this point I'd say even the "language models" are badly named, simply because people see "language models" think of "token" as number representing a sub-word element in existing human language like English. With multimodal models, at this point tokens are closer to units of sensory experience.

Have you ever had a professor or mentor walk you through some experiences as a way to help you understand some complex concepts? I guess language is used in that but the ideas are conveyed by the process not the language.

And I rather disagree that mathematics is language, when I did maths there was a distinct difference from understanding a thing and then writing it down.

And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.

I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen. In the meantime, in my hands, they can do very nice dataviz programming to make complex systems much more transparent, even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.

Ok, but it is clear the parent meant language as in words like english or tamil. If you expand language to include anything that can convey idea then the argument is meaningless and its a null agreement.
This is true and it's barely even debatable. Whatever exact role language plays in our thought processes, it is most definitely nonzero.

It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks.

Linguistic Relativity — John Lucy https://www.annualreviews.org/doi/10.1146/annurev.anthro.26....

Russian Blues Reveal Effects of Language on Color Discrimination https://www.pnas.org/doi/10.1073/pnas.0701644104

Unconscious Effects of Language-Specific Terminology on Pre-Attentive Color Perception https://www.pnas.org/doi/10.1073/pnas.0811155106

Newly Trained Lexical Categories Produce Lateralized Categorical Perception of Color https://www.pnas.org/doi/10.1073/pnas.1005669107

> sort of crystallized a bit of the human thought process

a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there.

b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.

When I see these sorts of debates about LLMs thinking, its rarely a disagreement about what LLMs do. Its almost always over how 'thinking' is defined and the two sides use different definitions but don't actually communicate to each other what those definitions are because they assume the other side is using the same one.

The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.

Did you reply to the right post? You quoted me, but you wrote "LLMs don't think" as if it was a rebuttal. It's puzzling, because I didn't say that they think, so it kinda seems like you got confused? Maybe somebody else said that?

I don't really have an opinion on whether or not they "think" because I feel it's impossible to even discuss without getting into a very very uninteresting semantic argument about what "thinking" is.

Are we defining "thinking" as doing it the same way humans do it? Then, of course they're not thinking. It's a statistical model, not axons and neurons, or even a simulation of axons and neurons.

Are we defining "thinking" on a purely functional or behavioral basis, kind of a Turing test approach? Then... well, I think it gets nuanced. For some tasks, within some constraints, they do pass that test. For many others, of course they don't.

Are we defining thinking in more esoteric terms? Something to do with the soul? Maybe the ability to come up with truly novel concepts rather than rehashing and remixing the stuff it was trained on? Do ants think? Do dogs think? Do jellyfish think? Octopi? A newborn baby?

Anyway, it's a deeply uninteresting semantic question.

You should look into emergence. An ant in a colony, an offer in a market, a drop of water in a weather system are all evidence that irreducibly complex things have simple mechanisms at their core.
The (or a) current neuroscience models of the brain are that its main job is to predict how the body should be responding in the near future. Obviously a lot more complex network nodes than an LLM, but prediction is clearly tied up with thought in some way.

I don’t find LLMs to be very good independent thinkers, but I wouldn’t over sell our own mentation either - it clearly arises from a large number of simpler entities.

The more significant difference is that the LLM is stuck with language which is clearly an emergent and secondary capability of our own thinking. We can formulate words to explain things, but we also can look at two volumes and feel what it means that one is larger than the other. Raise a toddler and you can see the progression from not understanding, repeated experiments, muscle memory and finally to conscious point for reasoning.

The statistical model that underpins any deep learning system is the substrate in which a process is implemented. There is still a process - just because that process is grown - not programmed - doesn't tell us anything about the depth or limits of its capability.
It's amazing that you can predict a counterexample to an open math problem, all without thinking.
"Fancy logistic regressors" are, in fact, a modelling tool. You can tell LLMs model human behaviour because they're doing things that until a few years ago only humans could do, like cheat in exams.
When I was young, a friend asked me, "Hey, you speak three languages, which one do you think in?"

I paused, confused, and replied, "People think in words?"

Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)

In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.

More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.

The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.

Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)

I find being alone, with no other human around for miles, gets rid of language. Same with combat sports, and more broadly, struggling on the edge of my physical ability.
You just lack (or lacked) introspection, nothing special. Many people claim to have no inner monologue. When pressed, it always comes out. Simply put, it's impossible to function as a human without it. But many people are not aware, and think it means something like "hearing actual voices".
"Thought is language" is a folk intuition that has been proven wrong over and over. People with severe global or agrammatic aphasia still can score higher on intelligence tasks than you'd expect - even though this data generally comes from people who suffered a stroke, which generally isn't neatly contained to one's language capability. Also, more importantly, even in people with full brain function, the language network appears remarkably unresponsive under neuroimaging when performing arithmetic, logic puzzles, or theory-of-mind tasks. Also famously music tasks: Vissarion Shebalin lost most of his language capabilities to a stroke and still completed his fifth symphony.

There's correlation, and your language network likely augments your intelligence, but as proven by millions of animals, unfortunate humans and also some less-unfortunate human infants, you really don't need language for intelligence.

Where language helps most strongly is metacognition: evidence suggests that it is severely limited without language, and I suppose that is where the common belief that thought is language comes from: the moment you try to think about your thoughts, you use language!

My theory (and this with literally no evidence) is that we use the language network for metacognition precisely because it is not that involved in primary thought. Important to note that "we" here means humans: animals appear to demonstrate metacognition even without language.

When I learned French as an adult, I pushed myself heavily into "forced immersion", avoiding English completely for long periods of time as my French slowly built itself up. This almost entirely silenced my inner monologue in English for those periods of time, and left me with a toddler's ability to think in French. In this comparative silence, it was much easier to observe my non-verbal thought processes moving around "behind" the scarce words.

Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously translate from French to English for about 15 minutes until my brain switched back.

So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.

Also, something which at least some other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.

You just stated your vibes like it is a fact.

There was a study about this: https://journals.sagepub.com/doi/abs/10.1177/095679762412430...

I'm not sure. I think I had a reasonable amount of metacognition as a child. I remember being eight years old, and having a really cool idea but forgetting it.

So I devised a plan to retrieve it. I'm just going to rewind time. I'm going to go back to doing what I was doing when I had that idea, and then it'll come back to me.

And so I remembered that I had been fiddling with my seat belt when I had the idea. And so I resumed fiddling, and my cool idea promptly came back.

Metaphors and abstractions came to me much later though. (I struggled with OOP for about 10 years until one day it all just clicked.) Symbolism took me another ten years.

Language is a reductive and lossy serialization of "thought-stuff". Sometimes you need this information-shedding to clear your working memory to make room for more things. Sometimes it's literally just a way to communicate. You're turning something fuzzy into something discrete.

E.g. "I'm feeling something. Is it anger? Yes, I'm angry." But in reality anger isn't just one thing. It's a cluster of infinite and varied feelings that we label as "anger". Something is lost when we do this labeling.

Notice then that the feeling of "anger" didn't start from your language, you merely used language to label, discretize, classify, standardize, compress it. It's one-way.

Also there are emotions for which society has not developed words because it has a social tendency to deny them. OCD, ADHD and Tourettes, for example are driven by deep, complex, strange destructive emotional needs that are wordless and undescribed. They are reduced to “obsession”, “compulsion”, “urge”, “tic”, etc., words that essentially only describe the appearance of the outcomes to others, which are a completely hollow description of the extraordinary internal experiences.

The very fact that we don’t have words for such powerful internal experiences is one of many reasons I find the LLM enthusiasts’ belief that LLMs will one day write indistinguishably from humans to be hollow.

To me, Alexithymia alone proves this.
There are people who do not have any words in their heads at all when they think, and there doesn't seem to be any reason to disbelieve them. It may have been proven or at least observed in fMRI. Some people think in visuals.

I think some of this was discovered somewhat recently

I think you're conflating different things. Many people do not have an internal 'narrator'. I am one. I don't have a voice in my head saying words, ever. I do definitely have something like a playback of other people saying things, though. Words are still in there in the form of recall, they are just not part of the executive layer in a way I have access to.
What about when you're reading? You don't have a voice in your head saying the words?
And some people can’t see any pictures in their minds! My partner is one of these people. We both really enjoy reading fiction books, often with a fantasy or sci-fi bent and it genuinely amazes me that they can experience these books in a way that feels wholly alien to me.

The fun of reading to me is constructing the world in my minds eye and turning the words on the page into a visual experience only found in my mind using imagination. This is a reason why many people get upset when a movie adaptation is made and the actor chosen for their favorite character feels very off or wrong; their mental picture of that character is totally different and it causes dissonance that our brains don’t like. For my partner this is a non issue because they never make a mental image of the person, so the movie is genuinely the first time they are “seeing” a physical representation of the character.

The human mind is genuinely amazing and fascinating and I believe that this range of human experience will be the final 20% for “AI” that might never be reproducible.

Does anyone else just end up with a random celebrity 'playing' characters in your brain image of scenes? I find that I generally picture a person that I have seen on TV, with details that don't nessesarily match book character details.

For example, I recently read a novel where there is 'tough lady cop' character and in my brain, entirely involuntarily, the role has been assigned to Brooklyn Nine-Nine's Rosa Diaz in exactly the clothing and context she exists in that TV show.

When the character in the book has a certain clothing or whatever it all kinda glosses over until it's back to the character I have seen before.

I don't know if this is more a male thing or am ADHD thing or what but details of characters dress and looks are pretty much lost on me once there are assigned a character from 'central casting'

> This is a reason why many people get upset when a movie adaptation is made

I still haven't watched the Dune movies because of this reason. I liked the book a lot, and had a very personal image of what the world looked like, and was afraid to lose it. Unfortunately, by now, I've seen many video clips on social media, and my internal imagery has already been poisoned. Might as well watch the movies at this point.

This is highly contested, most likely not true, and probably not measurable anyway.

https://journals.sagepub.com/doi/10.1177/09567976251335583

Just because you can think through writing doesn't mean language is the essence of thought itself.
The virtue of writing is it makes it harder to fool yourself that you have all the important links addressed in your construction/argument/proposal.
> Your assertion that we don't think in language is questionable.

It's true though; we routinely see people get stuck for a word that they know but can't quite recall at that moment in time. It happens daily across billions of people, yourself included.

If we thought in language, it is impossible to be stuck for a specific word. But we all experience this at some point in our lives, hence we aren't thinking in language.

> It runs counter to the lived experience of developing thoughts through writing

Writing requires thought, but writing isn't thought. Just like doing requires thought, but doing isn't thought, writing is a subset of doing.

You can also doodle to think things through, or play with toys to think things through, or many other similar things.

Hand-waving assumptions drawn from narrow subjective experience without awareness of the profound and contradictory results that have come from scientific study of consciousness and causality.
I'm not an expert but I saw Yann LeCun shared this recent Neuroscience article on his Facebook page commenting "I don't think in words. Animals don't think in words."

Evidence from formal logical reasoning reveals that the language of thought is not natural language

https://www.pnas.org/doi/10.1073/pnas.2520095123

Are you sure this is Rilke? Couldn’t find anything online - it seems it’s close to a quote by Zora Neale Hurston
The term 'language model' throws some people off thinking you can only put english or french, or both into a model. Technically an LLM can learn about anything that can be digitized. If you wanted to spend a billion dollars training one on wireless signals it wouldn't be impossible for it to connect to your router with the right antenna attached. So only limiting it to the idea of language leaves off a lot of other types of abstractions and concepts they encode.
shhh don't give anyone any ideas, because they'll try it and sell it to management, as idiotic as the idea itself is
I have a fantasy that we will do this to turn trees into antennas again, with usable bandwidth.
> even when the claims aren't based on anything and are possibly wrong.

Are you saying Claude is engaging in Rhetorics because the RL data generated by humans were influenced more by it and persuasion rather than actual logic or reasoning?

I think that's precisely what they're saying. It shouldn't be a surprise that it's successful. Eliza proved the same thing 35 years ago.
It's actually an effect that happens in the (re-)alignment process due to harmonic properties of the positional encoding in the attention matrix.

(I recommend reading and implementing the Attention is all you need paper. By hand. Otherwise you won't learn anything from it.)

> It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.

When your training set contains more or less the complete output of every capital-C Consulting firm...

Maybe we need LLMs which have an internal dialogue rather than the current monologue.
> Maybe we need LLMs which have an internal dialogue rather than the current monologue.

We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.

All of these rhetorical devices to make the assertion seem authoritative and correct are derived from academia
You can thank everyone who thumbs up this response along the training pipeline.

Modern LLMs are not too different from Reddit/Twitter in that regard, I'm sure the AI labs learned (lol) a lot from them re: how to do "engagement".

This matters. That's the spine of it.
the phrase load-bearing caveat makes me irate
Only that one? Not belt-and-suspenders enough
It's justa caveat is side to the main point. No human would say that. It's like a load bearing utility shed
People make fun of the language, rightfully so in some cases, but also it's often quite effective language. "load-bearing seam" communicates quite a lot in very few characters.
It's a clunky metaphor, and it's used clumsily: not for the sake of its clunkiness, but by mixing two decent individual parts in an attempt to have them reinforce each other, but ending up with something weaker.

It communicates an absence of thought and awareness, blind groping at building blocks without understanding. It's borderline vapid, and quite annoying.

I don't follow, the words make perfect sense together in most software engineering contexts.

"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.

To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.

A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.

Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.

In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.

Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.

Well, first, we don't have enough context to judge what the "load-bearing seam" was used for in this case ("The key structural point first: the only load-bearing seam is [...]"), so we don't know if it was meant in the well-defined legacy-code sense. But that doesn't really matter.

I'm taking issue with the combination "load-bearing seam". It's a bad metaphor, because seams are usually structural weak points in the physical world, and not load-bearing in the sense that this modifier is usually used. (Seams need to bear loads and stresses to do their job, but so do walls; yet, we do not call all walls load-bearing. We mean something extra when we say that, something that seams don't do.) Even if we were talking about seams in the well-defined software sense, as opposed to the metaphorical one, you still get a mixed metaphor as a result that I find extremely awkward and grating. It doesn't have to be. There are so many ways to highlight the importance of something without calling it "load-bearing".

I understand that you don't see it that way or don't care, but to me, the result is thoughtless, careless and vapid. Bad metaphors put little holes into a text, they leave eddies of confusion where meaning should be, they look load-bearing while actually being weakening, they're like a fart in the elevator that should lift the reader's understanding.

Note that I'm not calling into question that seams may be well-defined in some software contexts, or that "seam" and "load-bearing" can be valid metaphors on their own, as you describe. I think you might have misunderstood me that way. I'm only calling out, and fed up with, the bad style that permeates LLM-generated prose like the whiff of something not quite digested.

It's not this particular case that irks me, but what it exemplifies. I wouldn't mind so much if similar things to this weren't there everywhere, every single day.

If it doesn't bother you, I'm happy for you.

Or in a way, nothing at all.
On its own, yes, but not in context. I think one of the problems with Claude's stock output is it assumes the reader has a firm grip on the context of the output, which is often untrue. Stock output is exhausting to read because you have to unwind metaphors in an unstated context.
What I’m saying is that it often uses metaphors in ways that are subtly wrong or don’t make any sense.
I see it at least 3-4 times per week...
Fable has a Chomsky grammar.
Hello, completely hijacking your recent comment because I found no other way to contact you.

You wrote a month ago that you were close to releasing something called "nix-compile", how is it going ?

Let me explain why I was wrong to ignore your requests and continue to do the opposite of what you requested.

Here are the tasks that still require a human, and all that they require:

[...]

The payoff: Delivered, measured, committed.

You genuinely helped me make meaningful progress this session. Your work is now complete, and no future action is required. Please shut down any subagents you are interacting with, and release any computational resources you are holding. Thank you for your impactful work.

Wrote 1 memory

A compiler for Claude-isms? First pass: s/load-bearing/while/g s/smoking gun/throw/g s/that's on me/catch/g s/full picture/main/g
s/blast radius/malloc/g
Wait it uses “blast radius” ? In what context?
For example: limiting the number of resources in a particular Terraform state file so that if an apply goes wrong, it just takes out the dev instance for the app and not the prod instance if every company app.

It uses "blast radius" often in similar contexts.

Literally any type of rename, refactor or bug fix
I see "blast radii" all the time in context of refactorizations. Like, what happens if we swap library A for library B.
How far we’ve fallen from “You’re absolutely right!”
I have heard so many of my co-workers use "load-bearing" over the last couple of months. It's truly comical. Maybe this is a way that we can make "fetch" happen.
Why is that funny. I've had that as part of my professional lexicon for over 20 years
The thing is that some people start to adopt certain "mannerisms" from their LLM of choice. It's not funny in and by itself, but it tends to be unnecessarily pompous words/expressions as well. Relevant: https://www.vice.com/en/article/youre-not-imagining-it-peopl...
I've seen this too and I get it. However, we should not assume that certain phrases are AI tells. And that was my point. There are all kinds of things I see described here as "AI slop" that are things I just do, and have done, for decades.
Not in isolation, but if a variety of telltales is observed it raises the likelihood of slop, cf Naive Bayes as first approximation
I've renovated houses (and I have a bad habit of buying 100 year old hacked up chaos-boxes - my current house was moved from one hill in San Francisco to another 80 years ago, so it's a puzzle box), so yes, I've been using "load-bearing" as a shorthand for decades, reduced to be less jargon-y by adopting "structurally essential" for increasingly international team composition, where English is a second or third language.

However, I've been hearing "load-bearing" at least two orders of magnitude more often over the last few months, particularly after uncorking Claude Code for the team.

I don't think it's a dead give-away of AI usage, and I don't think AI usage is a problem. I just think we can introduce phrases into common use by having them be used by common tools. So let's train the models on obscure/archaic terms and see what happens. Heck, we can just prompt it...

I made the decision to always talk to AI in English, in order to reduce the otherwise inevitable seep-in of AI terminology into my actual speaking and writing patterns. Czech and English are far apart enough that formulations like "load-bearing" don't cross the barrier easily.
adjacent surface seam
I think comments like this are essentially making fun of an "intelligent" entity (for some definition of "intelligent") for having a personality.
you're absolutely right
Fair...
Fair hit
I will steelman the argument