|
|
|
|
|
by braebo
5 days ago
|
|
I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v... That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good. |
|
I felt better about Fireworks leadership after hearing some of them more candidly on podcasts (there are 7 founders!)