Hacker News new | ask | show | jobs
by Terr_ 1 hour ago
> The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true.

I think even that's too-optimistic: The model is a document-extender, so its baseline "truth" is whether a token seems like it would would fit-next in a partial document, based on prior documents. (This is usually not the kind of analytic truth we're interested in.)

So if we peek at vectors and weights, what we'll probably end up measuring is different, like the mood and tone of earnestness and conviction for whatever tokens are about to get emitted next. That may mean dialogue for a fictional AI Assistant character, a narrator, or an impersonal memo conclusion paragraph.

Is that the what we really want? Well, if our document described the character as Yoda, then The Force connecting all existence ends up truthy. Using "a really gullible person" can repeat anything you supply as truthy. Even if we set things up as "a respected encyclopedia article" or "a relentlessly logical super-genius"... we're not enforcing logic, but truthiness. [0]

In other words, we can't peek at ideas that require a kind of mind that isn't already there.

[0] https://en.wikipedia.org/wiki/Truthiness