Hacker News new | ask | show | jobs
by Aurornis 29 days ago
> It'll almost certainly be worth it, given the abusive behavior we've seen and will continue to see from the major closed-model providers.

The proper financial comparison for GLM-5.2 would be one of the providers on OpenRouter or renting a server as needed. Compare apples to apples.

You will almost certainly never break even compared to paying per token.

Local LLMs at this scale are only worth it if you have extremely strict requirements that data not leave the premises.

3 comments

Or if you want to hedge against the various tail risks of third-party providers raising prices or denying you service or somehow abusing your data...
> hedge against the various tail risks of third-party providers raising prices

They could 10X the prices and you’d still be better off. It’s also unlikely that prices go up enough to warrant a $100K local investment to prevent paying a couple bucks per million tokens.

> or denying you service

I guess you’re not familiar with OpenRouter? There are many providers there. There are providers outside of OpenRouter. There will always be someone to take your business.

> or somehow abusing your data...

If data security is your concern then you’re better renting a server as needed still.

If you cannot tolerate any data leaving, then local models are the only way. You pay a high premium for it!

People seem to miss that with local models you can have them burning their wee digital brains out 24/7, which is a different class of AI usage than that from online models even at a few dollars per million tokens.
There's a definite psychological branch point. With a remote provider, no matter how readily you can afford it, your mindset is always going to be, "I should think twice about what I'm doing. I hate to waste tokens." With your own hardware, your mindset is more like, "I should try to get more done. I hate to see this thing just sitting there idle."
Raising prices is not a tail risk, anything a local LLM setup can do for you can be done by any cloud provider, with the same capex as yours (or less), there is no moat here, so it is highy price competitive and will remain so. If you want to speculate on hardware shortages, that is a different business altogether and you need no janky garage setup to profit.
Also agreed, it's definitely a sucker's game to run a high-end model locally, by any objective measure.

Still... if it's not your weights, running on your box, you're always going to be behind somebody else's 8-ball. Everybody has to decide for themselves where their priorities lie.

Never say never. When the free money party stops, then those token costs are going to have to go up and up. The fact there’s such a glaring disparity between the cost of running AI locally and the pennies it costs to use an online model shows how heavily funded those platforms are right now. This is not and cannot be sustainable.
> When the free money party stops

The Openrouter providers the GP referenced were never at the "free money party". The actual cost of running something like GLM5.2 is well understood and tokens from those providers are not sold at a loss.

Obviously running things locally is more expensive but that all comes down to economies of scale. GLM5.2 is as expensive as it will ever be, barring an increase in demand that forces/allows providers to realise windfall gains disconnected from their underlying costs (always possible, but not the point).