Hacker News new | ask | show | jobs
by hodgehog11 12 days ago
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.