|
|
|
|
|
by yashchimata
13 days ago
|
|
One thing i keep thinking: you only run the pelican once per model. Run the same model a few times and you get some different pelicans, so some of "this one is better" might just be which run you picked for it. Would love to see 8 runs per model side by side. I bet for two close models, the gap between runs is about as big as the gap between the models. |
|
I built a whole ELO scoring mechanism a while back, described here: https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-...
I probably should spend some time on this now, even though the benchmark itself is feeling a bit stale. There's still a lot of demand for a gallery!