LLMs fail in such bizzare, obscure methods (to the average observer) at times. Sometimes even simple questions ("who was that x person who was super famous I'm thinking of") type questions fail terribly.
The more vague and non committal and hand-wavey and subjective the field for AI to answer, the better the results (imo).
The more vague and non committal and hand-wavey and subjective the field for AI to answer, the better the results (imo).