| Here is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception. If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer. |
>If you want to know this specific detail you might have to watch the movie yourself.
GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure.
GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer.
One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it.