| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by mierz00 134 days ago

If I am being honest, the value came from doing evals and testing against different models.

Essentially all I needed was a way to upload a data set, run tests against that data set and spit out a percentage of pass fail.

Braintrust makes this pretty easy, but If I was to do it again I would vibecode the same functionality.