Hacker News new | ask | show | jobs
by tsimionescu 9 days ago
Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them to give you bad textbook answers - but also don't give them AI for things they don't know, perhaps.

I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do.

2 comments

I've used LLMs to bootstrap successfully in a decent amount of things at this point.

Anyone trusting AI as the single authoritative source of information is stupid - but this follows from the fact that trusting anyone as a "singular point" as a source of information is stupid. You corroborate, you intervene on the world to test your mental model, you discuss with other people. That's what learning is. I've never learned from start to back to a textbook before as the single source of information (besides one philosophy of science textbook; in which I spent a month digging around adjacent fields, and then it just so happened that that one textbook synthesized every piece of information I looked up, and it was mostly a consolidating review).

If your study pre-supposes certain courses of action and artificially constrains the action space for the sake of "reproducibility", you may get a result, and a "scientifically rigorous one". But it's not going to say anything about reality in any meaningful way. While anecdotes and the complexity of real life isn't "science" (in that it's a controlled, repeatable, interventional experiment that's subject to a community of critics who want to hold you up to standards of rigor), there's far more truth in how people actually proceed and engage with these tools.

> but also don't give them AI for things they don't know, perhaps

The study doesn't show that at all. It didn't test actual AI.

They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.

Because this is exactly what they controlled for. FTA:

> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

(emphasis mine)

> precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool

That only works if they're experienced with model(s) of that level of unreliability and this is presented as one.

If they're used to a model that's more capable, and think the test model is similar, that's a huge confounding factor all by itself. It's not quite like giving fake credentials to a guy off the street and presenting them as an expert, but it's largely similar.

Just really really missing the point here in a kind of revealing way.. You probably have access to some really reliable models if I had to guess!
I've never used AI except for sometimes getting distracted by google search's builtin wrongness factory. So whatever "revealing" you think you found is completely imaginary. Rethink your assumptions here.

So please make an actual argument. How am I missing the point? It's true that someone trusting the AI in this test is not practicing "sensible delegation to a reliable tool". But what actually matters is whether they are practicing "sensible delegation" full stop. There's a big difference between "the subject inappropriately trusts AI in general" and "the specific test setup deceived the subjects". In the latter case, the attempt to remove the "sensible delegation" factor failed.

Edit: And any argument that uses "any LLM they would normally use is unreliable" as a basis is begging the question. If you can just assert that then you don't need to do anything to disprove sensible delegation. But if you can't assert it, the proof doesn't work right. So either the proof is pointless or it's insufficient.

The article, and the paper it references, is not an argument or "proof" of anything. This is not a theorem. The point is, given access to the same relative information as anything, there is (supposedly) a greater trust, there is less I-dont-knows with something in the form of a chatbot vs something else.

Whatever you're bias is or not here, the point you are missing is this is not about any given AI, or even about any given AIs "reliability" or not. It's not even, really, about "delegation" itself. It's just studying the supposed correlation here between uncertainty and one certain form of a tool.

There is no damning, sweeping thing to argue for here, this is not an editorial or an opinion, and does not purport to even be some big finding I would say. It's a pop sci article about a study done by (I presume) sociologists.

So yes, I would either way say you missed the point here.

Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

Is that what you're selling us?

So in 18 months, we'll just rinse and repeat?

They provide a sample of hallucinated answers from ChatGPT at the end of the study.