Hacker News new | ask | show | jobs
by peab 17 days ago
at this point, it's pretty easy to create evals/benchmarks, and then run the latest model on them.

LLMs are so easy to swap out, so having good benchmarks/evals are pretty useful.

Even then, a lot of the time the model improvements are so obvious that you don't even need an eval.