Hacker News new | ask | show | jobs
by juliangoetze 16 days ago
> This supposedly is better than KimiK2.7

How can you tell?

I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks.

Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?

1 comments

Most of the random comments you read on HN and reddit about how nice/bad various LLMs are, is basically based on the commentator's "vibe" about it, and almost nothing is grounded in evidence or actual usage. Don't read too much into it, want to know how good a model is? Run it with your own non-public benchmark, basically the only way to get proper answers you can somewhat rely on, everything else is manipulated, misunderstood or over-relied on.
I thought HN was different. And yeah, wherever I go, my timeline is full of Opus is so bad today and I will switch from Fable to 5.6 Sol, it's 1.5x better and vice versa.

Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.