|
|
|
|
|
by peab
17 days ago
|
|
at this point, it's pretty easy to create evals/benchmarks, and then run the latest model on them. LLMs are so easy to swap out, so having good benchmarks/evals are pretty useful. Even then, a lot of the time the model improvements are so obvious that you don't even need an eval. |
|