|
|
|
|
|
by hedora
33 days ago
|
|
In your box plots, 4.6 sonnet wins over all (even opus 4.6, the 4.8’s and fable). That’s not super surprising to me, but, given the apparent randomness of the stack ranking, is GLM actually worse than any of the Anthropic models? This looks like a 10-way tie to me. |
|
We now use Sonnet 4.6 for a number of internal use cases we wouldn't have considered otherwise.