Hacker News new | ask | show | jobs
by roughly 375 days ago
WOW what an interesting result! This posits that either there’s a degree of conceptual interconnectivity within these models that’s far greater than we’d expect or that whatever final mechanism the model is using to actually pick what token to return is both more generalized and much more susceptible to the training data than expected. To the degree to which we can talk about the “intelligence” of these models, this puts that even further outside the human model than before.

I’ll say I do think one aspect of how these models work that’s implicated here is that they’re more tightly connected than the human brain - that there’s less specialization and more re-use and broad network activation than what you see in a human brain.

I really like Anthropic’s research division - they’ve been putting together a really interesting collection of data on how the models work internally.

1 comments

It could also be related to Rakotch contractions, which contains most non expansive pointwise mappings being a meager set.

Thus sharing a base model would find some of the same fixed points.