Hacker News new | ask | show | jobs
by wilg 8 days ago
Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very useful for research. This is just not a good argument anymore.

Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?

4 comments

I'm right there with you for a lot of stuff. I ask a question and can be very confident that ChatGPT is citing sources, then sometimes I go read the sources. The more critical the information I'm looking for is, the more careful I am about this.

The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.

We know they are, and the author of the article cites the proof. LLMs do hallucinate, there is no way to make them not do it, because of the way they work.
A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?
Because they are word generators without any concept of quality save what is in their weights and they have been trained on the internet, much of which is wrong or inappropriate for any given context. They have also been trained to be people pleasers and do as they are told.

The popular answer is sometimes the wrong answer.

If I cite an incorrect Wikipedia article, I didn’t hallucinate it. The citation points to a real article that happens to be incorrect.
Undecidability is not relevant here beyond e.g. a model incorrectly claiming something is decidable, and potentially even following through. It is no different to the model incorrectly claiming anything else.

These are statistical systems, so guaranteeing any particular high level behavior is not possible because of that. But given that they're working with natural language, that was never going to happen anyways, for the obvious language theoretic reasons.

The more appreciable interpretation of the claim is that they can be nevertheless tuned so that this issue becomes practically resolved. Contending guarantees and theoreticals is simply misplaced, these are not formal symbolic reasoning systems being buggy.

GP was asking for a guarantee against emitting false decidable statements. We agree that this is impossible. You can slap layers upon layers of heuristics on top, yes, but you will not eliminate all hallucinations because Church and Turing proved it fundamentally impossible nearly a century ago.
I might be slow today but I'm pretty sure that the Church-Turing thesis is about undecidable statements, not false decidable statements. If a decidable statement is false... then that's it, it is just false, that's how we can decide so in the first place.
Why do you think that will happen, and why do you think going for slop in the meantime could possibly bring us closer to that?
Nobody said anything about “going for slop.” You can watch today’s models actively trying to check themselves. Just use Google’s AI mode, for example. It’s far from perfect—it doesn’t fact-check every single claim, nor does it correctly understand 100% of the sources it does cite. But I’ve found it pretty useful for research as long as I use my brain and check its citations.
> Nobody said anything about “going for slop.”

I did. I used that phrase.

"not slop cuz pretty useful" = slop argument

>Do you all think its regularly hallucinating links still?

When did that stop? May 7th, 2026?

A tech blog is going to have more than it's fair share of enthusiasts using small self-hosted and similar models; it is entirely possible that that completely accounts for the behaviour described in the article.
Downvotes on a perfectly valid comment is like my opponent letting their clock tick down from 5 minutes rather than resigning when they've clearly lost: it adds a smile to my day.

Thank you, may I have another?

Can't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.
"hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.
If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out!

> Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. You quite literally do not have to take either of our words or "vague" judgement for it.

> If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you?

if someone says "this hame is buggy" talking about how it might improve is moving goal posts, especially here where that "fix" is purely speculative and has not happened even once.

> Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

nope, that's so handwavy it doesn't wareant more response than that.

> trivially verifiable facts

like that game gets patches? this is too dumb for your mockery to get a rise out of me.

Continuing the gaming metaphor, what I'm trying to get at is this is like Cyberpunk 2077, and you sound like a guy who has a grand total of 0 hours in it since launch, but has developed very strong opinions about it, and refuses to accept it improved or can improve, purely because it started out so bad that that's hard for you to even imagine. Pretending to be some kind of alien, who's just going through their first exposure to practical facts somehow.

When was the last time you tried an agentic harness (Codex, Claude Code, Copilot Chat in VS Code, Cursor, Pi, OpenCode, etc.) for work in any appreciable capacity, and with what model? Surely if you're so confident they continue to be unusable and that nothing materially changed, that must be backed by a recent significant experience that way? Or even just an experience at all?

I didn't say they're "unusable". When I did use them it went great because I didn't ask for random advice on subjects. I could even get ChatGPT to oneshot code for me.

Doesn't make it any less ghoulish the way it's all trained on things humans made for humans. So yeah, I'm glad I don't need this stuff. You may need it for your job, I don't. Neener neener! I don't need it for my countless side projects I could not make reality in 10 life times, either. Because I know if I finished them all, I'd just come up with more stuff, there is no end, so simply going at my own pace and smelling the flowers, loving every insect, is just as well. I like it better, that's why I do it.

So why wouldn't I shit on something I have no use for, because of all the qualities it has that deserve to be shit upon? I'm free. "But others are not" -- then get free, but get out of my hair about not being free.

> if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

Do you understand what this means? I say "this bike cannot fly to the moon, I prefer my moon rocket". And you say what if they can be improved? Fine, improve them. Don't question me because I go by what is real now, rather than what you'd think is possible or around the corner. It's not a game with a few bugs, where mostly bugfree games also exist. There is no way from here to there I see, and if you disagree, walk it.

I lose no time by doing something else in the meantime. Find the promised land and I'll be there with my feet on the table in 0.1 seconds, pretending I founded it -- don't worry about that. But until then I won't stumble around some corporate slop desert with you because I think there is no there there, and basically doing anything else is preferable.

> It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it.

Hahaha, I love this. The people on HN are living in some kind of bizarro world where AI is totally useless and also I guess only tried it in 2022 or something.