Hacker News new | ask | show | jobs
by mppm 38 days ago
This is cute, but here is the result I got by always clicking the longest answer, or the first one if two seemed equal:

Scientific Estimate: 71,650 words

"Unbelievable. Are you actually Stephen Fry in disguise?"

Core Basics: 16/20

Intermediate: 15/20

Advanced: 19/20

Expert: 18/20

Grandmaster: 16/20

This is significant beyond this particular app, because biases like this are found all over the place in popular LLM benchmarks.