Hacker News new | ask | show | jobs
by llmslave 20 days ago
These models are never as good, the benchmarks dont tell the full story
3 comments

The reality is most people building their own models and providing that alongside SOTA ones don't really care about how great these models are. They just prove that 'hey we are smart enough to build our own models so you can trust us instead of going with a single provider like Claude via Claude Code', also a cheap alternative for cost sensitive/free users - at least this was the case for Windsurf, not sure if Devin Desktop still has that tier. They just need to hillclimb the benchmarks and show something reasonable enough there.
Benchmarks are just vibes with error bars... wake me up when it survives a week on a real codebase without hallucinating a package that doesn't exist.
Funny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.
It's simple: close weights = not welcome.
It's almost as if HN users aren't all the same.
all the open source models are a waste of time relative to the bleeding edge from openai/anthropic
Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
I use GLM or DS4 to help me draft a better initial prompt with more information that I then give to Sonnet 5/Fable/GPT5.5. While benchmarks show the open models close to frontier level, my experience with them is drastically different. I have high confidence that Fable or GPT will 1 shot solutions.

At least with low level programming languages. They're all very good for webdev stuff.

yeah but why waste your time on these models, just use the one that gets the better results
I actively prefer GLM-5.2 for some tasks. For simple tasks the results are just as good as e.g. Opus, and it produces results significantly faster.
Because you can get them from more trustworthy providers or with hardware encryption.
i trust anthropic/openai with my data far more than some random startup.
I was going to respond until I saw your account name lol.
haha i outsource my thinking to the smartest model
At work I wouldn't want to use anything else. Compared to my salary a Claude subscription (or two) is cheap

For hobby projects I've completely switched to DeepSeek v4 pro. I spend less than on a $10 Claude plan and am not subjected to quota limits (when I have time and motivation, the last thing I want is a 5 hour quota running out). And the difference in model performance is fine for those smaller projects, most of which will end up abandoned or in a state of "good enough" anyways

And for utility tasks, those 30b models are also great. I'm a big fan of gemma4

ive just got better things to do with my life than fuss with an inferior model. its like why hire a dumb employee over a smart one
I think you misspelled "I've got plenty of money".
200 bucks a month?