Hacker News new | ask | show | jobs
by dom96 16 days ago
I love LLMs too, but I am concerned about their cost. They are all still very subsidised. Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?
6 comments

> They are all still very subsidised.

I think the opposite: I think the frontier labs have good margins on their inference unit costs.

We can already see what it costs to run near frontier-size models. There are independent business pivoting to serving these models at reasonable prices and they're competing on OpenRouter for costs much lower than frontier labs.

> Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?

I would bet good money on prices going down significantly, not up.

If we get to the point where you can run an Opus 4.8 model on your local computer, it's going to be even cheaper for a datacenter to serve it on their hardware. That means prices crash, not that they're going to rise.

They may have good margins, but a few things are still true:

1. Much of those profits have to be immediately reinvested into model training runs to avoid being lapped by competitions.

2. Unit costs are irrelevant when the labs don't price per unit, and instead charge, for instance, $200 / month for $10k worth of tokens.

This isn't a steady state. Whatever the current situation is, I doubt it's sustainable.

> 2. Unit costs are irrelevant when the labs don't price per unit, and instead charge, for instance, $200 / month for $10k worth of tokens.

Cost to generate all of the tokens divided by revenue generated by selling those tokens is what matters.

The subscription plans confuse a lot of people because that's what they see. They're not seeing the gigantic API bills from all of the tokens going into enterprise use cases.

The subscription plans are a small part of their income. Most users aren't maxing out 100% of their plan usage every week. I wouldn't be surprised if their average plan user was using less than 50% of their monthly quota each month.

Plans like that can produce a net increase in profit if they get consumers interested in the brand and pitching it at work. Giving them some extra token headroom on their $20/month or $100/month home plan is money well spent if it gets all of a company's developers advocating for enterprise plans with budgets exceeding $1000 per person.

enterprises are not dumb, they look at the cost of their ai investment and reevaluate it every quarter. Uber recently capped their AI spending per employee, and then there's this article a couple of days ago: https://finance.yahoo.com/technology/ai/articles/ceos-being-...
You’re all wrong and effing stupid

Until you bring reinvestment into your analysis stop posing

The subscription based plans are heavily subsidized, but the direct API inference pricing (which larger companies need to pay) is profitable.

Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.

> Using a full Claude Max 20x plan to 100% of weekly usage

I doubt many of their customers are on the 20X plan. Of those, I doubt many of them are using 100% of their weekly usage regularly.

Comparing the 100% maximum usage scenario of their most discounted plan against the API cost has been a trap in this conversation since it came out. I bet if we saw their financials it would be a tiny sliver in a pie chart somewhere.

True, it'd be a whole other situation if the tokens limits were cumulative. I guess it would all come down to whether their Claude Code subscription plans are turning in a profit or not.

At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.

https://arstechnica.com/ai/2026/04/anthropic-tested-removing...

$2k through the API

Which costs them $200 to serve.

Yes, it is a 10x markup on the API prices. Depending on whether you factor in cooling costs, data center staff, etc. Or GPU costs and the electricity the GPUs are using only.

Either way, inference is very much where the money is made, training is where the money is lost.

I thought hardware prices would always just keep going down.
That is a great comparison. The problem is when costs become prohibitive for new entrants.
Interestingly enough, geohot also has an article covering this: https://geohot.github.io//blog/jekyll/update/2026/06/18/pric...
That's commentary on company valuations.

Token prices are going down. Competition is global. A company could choose to keep their API prices high, but if another company comes in at 1/10th the price for 95% of the performance then they won't have many customers.

You’re right, my bad, I read that too quickly
> I think the opposite: I think the frontier labs have good margins on their inference unit costs.

Interesting. Good enough to make up for training costs?

Also, where can I read more about this?

A few days back there was a post saying the only ones making money with AI are the ones selling the hardware
I mean, Pfizer has a good margin on every pill they produce too.
You can maybe run a local Sonnet-4.5-ish-level model (sort of) for less than the price of a new car, even at current massively inflated prices for fast RAM. This is probably not what you were looking for. But it's there. You could share one server between multiple developers. Maybe make a little AI co-op or something, with a pair of RTX Pro 6000 cards?

Also, DeepSeek V4 Pro is cheap via any commodity API, and DeepSeek V4 Flash is essentially free at API prices like $0.09/M, $0.18/M out. This is generally not subsidized.

For a more practical local setup, Qwen3.6 27B on a used Nvidia 3090 (US$1300) or two is surprisingly nice. It needs clear instructions and you can't use it for hands-off vibecoding, but it's actually quite reasonable for hands-on programmers.

I’ve got a pair of those cards and DS V4F is incredibly good. I’m happy I did what I did because I like this stuff but if you just want stuff then you are absolutely better off not spending $20k on two of these cards and using the API. This guy is absolutely correct.
Guarantee is too strong a thing to seek, but healthy competition makes it highly likely that the supply/demand curve will meet at a healthy place.

You're always guaranteed that you can stash away the open models!

We're supposedly getting Mac Studio with 1.5Tb RAM in 2 years. That would be enough to run an Opus-level model.

Of course, it will also probably cost somewhere around $50k...

But if local AI really does become pervasive, maybe it'll be one of the things people buy on credit, like cars.

"Of course, it will also probably cost somewhere around $50k..."

Whats to stop people remotely accesssing this? People already do this when working remotely in finance - they connect to a virtual environment that does their work in spreadsheets lmao. nobody cares about the lag, managers certainly dont care about sub-ordinates complaining about it - the same way nobody will care about a slight loss of quality if the economics make sense. frontier labs are screwed really.

Currently, because of the subsidies from the frontier models, demand is mostly for higher intelligence.

If subsidies do end, demand for price efficiency per unit of intelligence will go way up.And because there's many players in the market, this demand should be met by at least some of them.

would i hire a phd grad if they cost the same as a degree holder? sure, why not. but if they cost twice as much?
GLM-5.2 is runnable and downloadable today on a MacBook studio that costs a stupid amount of money. No one can take that away from you except by force though, if you want to set it up today.