Hacker News new | ask | show | jobs
by steveBK123 27 days ago
To reply to myself here..

I can STILL replicate this behavior in Google AI summaries 10% of the time:

"is <SOMEPLANT> ok for cats"

to which it replies: "Yes, <SOMEPLANT LONG SCIENTIFIC NAME VERBOSE PHRASING> is toxic for cats"

The other one going around this weekend: "how long hot dogs on grill"

Summary: "The hot dogs on your grill are likely around 5-6 inches long .. "

So scale this category of error to unsupervised agents with access to your credit card.

1 comments

This repros nearly 100% of the time on most LLMs, even the most advanced ones: https://share.gemini.google/u9NwYu7lbgxe
n=1 but I gave this to Sonnet 5 medium effort (free model) and it had no trouble with it
Try it without "reasoning". As you can see in my example (and GP), it meanders to correctness eventually after emphatically being wrong, and most reasoning modes hide that from you.

If LLMs worked the way people want to believe they do, there’d be no reason to start in the wrong place — a computer should have the facts!