Hacker News new | ask | show | jobs
by nijave 20 days ago
I've done a couple side by sides on web chat with the same prompt on Opus 4.6, 4.7, and 4.8 and the output gets longer/more verbose on version increment. The enerr variants are definitely much wordier.

On the other hand, the newer variants also tend to benchmark higher so it's not quite a clean argument of "hey the new version eats more tokens"

2 comments

I think both things can be true: new models benchmark higher and eat more tokens.
From my experience new models are slower and use more tokens even on questions which gpt 4 answered correctly. It is mostly because newer models tend to be more verbose (even with prompt requesting short answers).
Unless somebody improved on the underlying transformer architecture... Surely AI is smart enough to do it by now
I've done a couple side by sides on web chat with the same prompt on local 4b, 14b, 32b open models and the output gets longer/more verbose on version increment.

Its rather frustrating, slower tokens and more tokens.