Hacker News new | ask | show | jobs
by jacobgold 29 days ago
> "~$40k At this price level, you get the next step up in model intelligence. Something pretty close to Claude Opus."

That is equivalent to 16.8 years of Claude Opus 4.8 or Codex GPT 5.5 at $200/mo.

I'm a huge fan of running local models, but they're still wildly expensive, lower quality, and possibly dangerous (if backdoored). I sincerely wish this wasn't the case.

5 comments

That $200/month is already more like $4,000/month if you have to pay full API pricing - "enterprise" companies for example. That drops the equivalent to 10 months.

(I'd be surprised if that local rig really can drive the equivalent of $4,000/month of API spend though, given that a local rig can run prompts in parallel a lot less effectively than Anthropic's many data centers.)

I think the decode phase of inference typically uses local compute resources poorly due to the very small batch size. If you can run many inference tasks in parallel, this will make local inference more competitive to centralized inference, not less.
I agree with your point, but it should be noted that this assumes consistent prices for LLMs. The OpenAIs and Anthropics of this world are still selling the plans at a subsidised prices with the power of VCs, who are going to want that return some time.
VCs need to sit on ice for a few more years if they dont want it to pop
You can use a lot more tokens on hardware than you can spend on a $200/m plan.

Inwrnt through 1B tokens my first month with an OEM spark. That's more than $1k of opus. Not a fair comparison, because token patterns are different, but since that time I have also seen a 2-3x improvement in then speeds.from improvements in vllm (mainly MTP). DiffusionGemma is around 4x regular gemma.

None of the leading models are backdoored, that's nonsense. I've still never heard of a single backdoored model, and if one was found, it would be quickly eradicated from HF. This is a non issue.
Stop trying to run them locally, folks.

You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset?

Rent cloud GPUs!

You get to participate in the ownership, data control, price control, and hacking culture without having to Frankenstein some hobbyist box that costs a ton, is distilled down to functional uselessness, and is a PITA to maintain.

If I'm gonna rent cloud GPUs I might as well just use a subsidized cloud agent like Claude or Codex. As for depreciation, that is true, but the bet is that models get better for a certain parameter count faster than your hardware becomes obsolete, such as Gemma models for example at the same 30 billion parameter count being much better than some years ago.
This comment is like the antithesis of hacker culture. Can’t tell if being ironic or not.
Lobotomized RTX models are playthings.

People building this stuff are "year of linux on desktop"ing open weights AI. It's a huge opportunity cost - not just for you, but for the open source community at large.

You need to double down on big fat honking models that take multiple H200s to run. That's where the real power lies, and that's where our entire community needs to focus our efforts if we want to keep the delta between frontier and the proletariat small.

The more we build for people and enterprises to run big weights in private clouds, the better. That's the real treat to Google, Anthropic, and OpenAI. Your RTX cards don't make a dent in the death star.

You could be an IBM executive writing about the Apple I.
There's a world of difference between the Apple I and setting up CUDA drivers and Python.

You're in a small community of hobbyists. Cheap hobbyists who mostly don't pay for the stuff.

It's a bad growth market. It's a bad space to develop products. And it's as distracting to brilliant minds as bitcoin.

It's a suboptimal nerd snipe, and all the effort spent there is effort not being spent building actual frontier capabilities in the open.

> Stop trying to run them locally, folks.

"Locally" is a relative qualifier if one defines locality as not being reliant on a SaaS vendor. IOW, locality does not necessarily imply execution on machines specifically owned/operated by an organization.

> Rent cloud GPUs!

This would qualify as "locally" in the above definition. There is also a case to be made that h/w ownership (GPUs included) and operation can result in a net cost reduction for some use-cases.

However, where exposing intellectual property results in regulatory violations and/or undue legal exposure, running models "locally" is not only a good option, it is the only option.

> You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset?

Single mode fiber can serve for tens of years without problems and push the fastest speeds available today. I do not understand this comparison.

We don't need to own the hardware.

We need to own the software and the models.

Playing around with local models is like playing around with Ubuntu and Arch in the 00's. It's a fun toy, but it doesn't make a big economic dent, and it doesn't ensure we retain our rights and a slim capability gap against the frontier.

Developing software that works with big models, showing up with economic demand - that ensures that capability gets built and that open whittles away at closed at the very frontier.

More customers going to tiny hobbyist models also sucks oxygen out of the room for more large scale open models. We need to put economic demand on the larger open weights.

I see two futures ahead:

Future 1 - Big companies alone have access to the most productive models. Consumers play with API offerings and tiny RTX-scale models that lack the same capability.

Future 2 - A robust assortment of open weights models keep a very slim capability gap against the most mature frontier models. There's a viable economy around using and supporting these big models. Prosumers and enterprise can easily rent spot instances and spin up weights on-demand for a variety of tasks. There are rich model and fine tune marketplaces, a wide assortment of tools that can call these models, and easy tools to train models for any task from any foundation pretrain starting point.

I'd rather we go down path #2.

> You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset?

Like a car? Because I don’t want to depend on uber or a taxi service.

What does fiber have to do with anything? I don’t need the internet to run my local models.

Yes, "Like a car". LOL. You realise that many people in Europe and Asia do not own a car at all? Public transport, eBike / scooter, Tuk-Tuk, walk.

The local LLM "privacy" war had been already lost.

You realise how many people in Europe do own a car?

And how exactly has the privacy war been lost?

The privacy war has been lost in two ways (at least) 1) Running locally lobotomised models makes no sense; 2) as someone said here, the Gov will declare local AI a felony. And they will enforce it. So those "many people" will buy V8 cars limited to V4 and declared illegal to drive without registration and license, even locally in your own yard, and they may go to jail if they attempt to activate other 4 cylinders. Oh, wait ... isn't it how car laws work now? ;-)
Ludicrously paranoid take.