Hacker News new | ask | show | jobs
by slibhb 9 days ago
> but also don't give them AI for things they don't know, perhaps

The study doesn't show that at all. It didn't test actual AI.

They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.

3 comments

Because this is exactly what they controlled for. FTA:

> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

(emphasis mine)

> precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool

That only works if they're experienced with model(s) of that level of unreliability and this is presented as one.

If they're used to a model that's more capable, and think the test model is similar, that's a huge confounding factor all by itself. It's not quite like giving fake credentials to a guy off the street and presenting them as an expert, but it's largely similar.

Just really really missing the point here in a kind of revealing way.. You probably have access to some really reliable models if I had to guess!
I've never used AI except for sometimes getting distracted by google search's builtin wrongness factory. So whatever "revealing" you think you found is completely imaginary. Rethink your assumptions here.

So please make an actual argument. How am I missing the point? It's true that someone trusting the AI in this test is not practicing "sensible delegation to a reliable tool". But what actually matters is whether they are practicing "sensible delegation" full stop. There's a big difference between "the subject inappropriately trusts AI in general" and "the specific test setup deceived the subjects". In the latter case, the attempt to remove the "sensible delegation" factor failed.

Edit: And any argument that uses "any LLM they would normally use is unreliable" as a basis is begging the question. If you can just assert that then you don't need to do anything to disprove sensible delegation. But if you can't assert it, the proof doesn't work right. So either the proof is pointless or it's insufficient.

The article, and the paper it references, is not an argument or "proof" of anything. This is not a theorem. The point is, given access to the same relative information as anything, there is (supposedly) a greater trust, there is less I-dont-knows with something in the form of a chatbot vs something else.

Whatever you're bias is or not here, the point you are missing is this is not about any given AI, or even about any given AIs "reliability" or not. It's not even, really, about "delegation" itself. It's just studying the supposed correlation here between uncertainty and one certain form of a tool.

There is no damning, sweeping thing to argue for here, this is not an editorial or an opinion, and does not purport to even be some big finding I would say. It's a pop sci article about a study done by (I presume) sociologists.

So yes, I would either way say you missed the point here.

> The article, and the paper it references, is not an argument or "proof" of anything. This is not a theorem. The point is, given access to the same relative information as anything, there is (supposedly) a greater trust, there is less I-dont-knows with something in the form of a chatbot vs something else.

...Are you saying there is a difference between an "argument" and a "point"? And you accuse me of missing what people are saying..

Okay, they were making a point about how people delegate. They wanted to remove a confounding factor "so any reduction in judgment could not be explained as sensible delegation to a reliable tool." But because of how people judge things, they failed to remove that factor, and possibly made it worse.

> It's not even, really, about "delegation" itself. It's just studying the supposed correlation here between uncertainty and one certain form of a tool.

But they decided they cared about removing the "sensible delegation" explanation. I'm not imposing on that on them. They thought it was important to remove, and they did something that doesn't remove it at all.

Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

Is that what you're selling us?

So in 18 months, we'll just rinse and repeat?

They provide a sample of hallucinated answers from ChatGPT at the end of the study.