The reality is most people building their own models and providing that alongside SOTA ones don't really care about how great these models are. They just prove that 'hey we are smart enough to build our own models so you can trust us instead of going with a single provider like Claude via Claude Code', also a cheap alternative for cost sensitive/free users - at least this was the case for Windsurf, not sure if Devin Desktop still has that tier. They just need to hillclimb the benchmarks and show something reasonable enough there.
Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
I use GLM or DS4 to help me draft a better initial prompt with more information that I then give to Sonnet 5/Fable/GPT5.5. While benchmarks show the open models close to frontier level, my experience with them is drastically different. I have high confidence that Fable or GPT will 1 shot solutions.
At least with low level programming languages. They're all very good for webdev stuff.
At work I wouldn't want to use anything else. Compared to my salary a Claude subscription (or two) is cheap
For hobby projects I've completely switched to DeepSeek v4 pro. I spend less than on a $10 Claude plan and am not subjected to quota limits (when I have time and motivation, the last thing I want is a 5 hour quota running out). And the difference in model performance is fine for those smaller projects, most of which will end up abandoned or in a state of "good enough" anyways
And for utility tasks, those 30b models are also great.
I'm a big fan of gemma4