Hacker News new | ask | show | jobs
by bigyabai 37 days ago
> From GLM 4.7 flash

GLM 4.7 Flash is a 30b model that was far behind SOTA at launch, and I know that because I pay for z.ai inference and have run the model locally. Qwen and Deepseek V4 Flash have the same issue, and beg the question; are you really going to process a 64k agentic context at 450tok/s? That's 2+ minutes that you spend waiting for the first token to generate! Of course nobody can sell that as competitive inference, and it only gets worse with larger models. We're talking about non-interactive speeds, here.

If you're satisfied with small local models, more power to you. It puts you in the same barrel as Strix Halo enthusiasts or the guys that bought 2x3090s on Reddit. You are completely ignoring the market if you think that any of those SOCs are unprecedented or unparalleled for inference workloads, though. The free DS4 API is faster at prefill and decode, you could not give away Mac inference at zero cost and compete with what China provides for free. That's how far behind Macs are for local inference, to put things into perspective.

2 comments

I think you’re confused - nobody running local models is concerned about SOTA - that’s just marketing hype from large providers - we are interested in data governance, security, control and freedom. You can’t compare hosted services to local inference, these are two very different things, out of principle we aren’t interested in handing our code bases over to untrustworthy third parties.

On your first point, nobody is pasting 64k tokens at once as context, if you are you’ll experience very similar wait times even with hosted providers - context is built up piecemeal and by virtue of being context does not need to be replaced constantly - this is how all agents work.

100% local models are not SOTA, but they are good enough to be incredibly useful if you are a skilled engineer, and I understand the industry is pushing engineers to offload more and more of their work to incentivize higher token spending, but talented engineers can absolutely be just as productive using local models today - it’s just a different style of working that folks who have become accustomed to large providers can’t really comprehend at this point. They’ve vendor-locked themselves into a delusion that they absolutely need SOTA for everything and as a result see everything as black and white.

You sound like IBM in the mainframe era...