Hacker News new | ask | show | jobs
by csomar 14 days ago
You picked a bad example because this is a thing where LLMs excel. A human doesn't know that tomato doesn't play well in a fruit salad until he learns it; not that different from LLMs.

What LLMs lack, from my experience, is two things: First, a long term memory (even a zoomed out one) and Second, the ability to execute a multi-step task without fizzling out. Upon adversity, LLMs breakdown and start making shit up. The newer models are better (and that's the difference between Fable/5.6 and GLM 5.2) but still have a very limited ability to execute any tedious planning tasks.