Hacker News new | ask | show | jobs
by dozerly 17 days ago
I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.
3 comments

If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!
That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
It's not a benchmark, it's a meme
To be fair, the "they're trained on this benchmark" response is also a meme.