|
|
|
|
|
by embedding-shape
16 days ago
|
|
Most of the random comments you read on HN and reddit about how nice/bad various LLMs are, is basically based on the commentator's "vibe" about it, and almost nothing is grounded in evidence or actual usage. Don't read too much into it, want to know how good a model is? Run it with your own non-public benchmark, basically the only way to get proper answers you can somewhat rely on, everything else is manipulated, misunderstood or over-relied on. |
|
Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.