Hacker News new | ask | show | jobs
by thoughtpeddler 21 days ago
Anyone else have tips for how to build skepticism around this type of paper? I find myself for whatever reason more readily inclined to believe the Anthropic mech interp team's claims, but then after reading skeptical takes, I 'snap out of it' and more clearly see the still-unsettled science of it all, but I wish I had better priors. Although I follow this space fairly closely (versus the "average person"), I still feel under-equipped when facing research that might be equal parts marketing and science.
2 comments

For papers that make grand claims, they usually have poignant limitations. In this paper, it's clear that the conclusion is extrapolated from a small set of observable information. That's usually the recipe for poor conclusions, and they are up front about where their findings sit.

That's not to say these findings aren't valuable though. Their blog post summary is just a bit more hype inducing than the underlying paper.

The way I feel about this is that the emergent properties of LLMs seem to reflect our own human faculties. Whether that's because we built them in our image or whether this is a common mode of consciousness is definitely outside the scope of this paper.

garlic_enjoyer has already said valuable stuff, but you must realize that skeptical=/true. A lot of people simply don't know what they are talking about. I remember on one of the previous mech interp papers arguing with someone who just didn't even understand what the paper was saying and the experiments they had set up and so a lot of misunderstandings and wrong conclusions spilled from there. And it's kind of funny because you would certainly think he knew what he/she was talking about from how self assured it all was.