Hacker News new | ask | show | jobs
by zkmon 5 days ago
Why is it hard to come up with tests that closely resemble real-world usecases where the models are being used?
2 comments

It's not that hard - there are plenty of those.

They're mostly not very funny though.

Would it be any more applicable?

We see the same problem in database benchmarking.

The good thing about the deliberately-non-real world pelican case is that it gives a general impression of how much the model is improving because it's not likely that it's being specifically targeted at it, rather than a 'real world case' which might have been specifically optimised for.