|
|
|
|
|
by meander_water
7 days ago
|
|
The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development. The only purpose these metrics serve is bragging rights for the model companies. |
|
The model did fine.
Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.
In this moment, andai was enlightened.