Hacker News new | ask | show | jobs
by persedes 6 days ago
Glad to see a benchmark outside of programming tasks. Presumably this one was not part of the training and shows the models performance (or lack thereof) on knowledge tasks.