Just a funny observation. Every time someone proclaims "LLMs can't do X", a bigger, badder LLM that can in fact do X shows up shortly thereafter.
Clearly, Fable 5 didn't even have the decency to wait until the next model refresh cycle to show up. It was already sitting there waiting.
Either the capability gains in bigger, badder models are actually unrelated to "gotchas" being discovered, or LLMs are already acquiring Skynet levels of disrespect for cause and effect.
Or just AI denialists like to say "LLMs can't do X" even though they can and have been doing it for the past few months or more. They only get called out once the current SOTA LLMs get so good at it, that any rando can trivially and reliably falsify the claim on the spot with whatever SOTA LLM surface they have handy.
Which I suspect is what happened here, given the trail of smaller / local models that successfully answer the question, too.
That said, "curse of reversability" is real, as much for LLMs as it is for people.
I don't think it's solved this fundamental architectural problem by itself, it will have just squeezed the edge cases thinner. It keeps happening, people find a question it gets stupidly wrong, the vendors proclaim they've fixed it, then another one gets found.
What does it say to the second question? I've found Claude is one of the worst models with regards to pop culture knowledge like this, even compared to the Chinese open ones. Just curious, not really relevant to the initial post but I don't pay for it so I only have access to Sonnet.
https://claude.ai/share/5e7e09b2-a75a-4024-b261-9a1a4e063a8b this is mostly hilariously wrong. wrong tie colors, they did not replace their bassist with a drummer, two completely made up albums, the rob cantor song it is thinking of is "shia labeouf", and a few fan behaviours i think it just made up
For most cases probably not. It's just something I like testing new models on sometimes, the pelican riding a bicycle benchmark probably isn't that useful either.
Thank you for prompting me to try with my own obscure question, Fable was able to find something from a poor description. I've been looking manually and over many sessions with different models as they improve, none have been able to find what I was after. Fable 5 just one-shot the answer.
Every time someone somewhere says "an LLM can't do this", the next generation of LLMs gains one more parameter. Until that LLM can, in fact, do this.