|
|
|
|
|
by fastball
14 days ago
|
|
Yes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models much less effective: I don't want a model I use for coding to burn a bunch of reasoning (read: time) on trivial tasks. |
|
For example, Kimi 2.7 has been really effective for me despite having verbose thinking blocks, simply because it runs so fast. Speed-wise, it feels about like Sonnet, possibly faster.