Hacker News new | ask | show | jobs
by etskinner 14 days ago
Inference is cheap, only the training is expensive. Both Bernie and Trump have suggested doing a partial government takeover of the big AI companies to start a sovereign wealth fund.

So it would look like the government taking ownership, letting investors lose their stake, and then operating as inference-only, which would turn a profit

3 comments

> Inference is cheap

We have absolutely 0 hard proof of this. We have a lot of wishful thinking but no hard numbers, audited numbers from any public entity.

I'd love to see them if they are available.

Where have you looked? OpenRouter? Your own experiments? From running various models locally on my MacBook, and paying for the laptop and the electricity to power it, but not the training run, as all I did was install some software that downloaded models from Hugging Face, yes it's cheap. Well, the hardware was several thousand dollars, so not cheap on a personal level, but not unaffordable either.
> OpenRouter

Do we have the balance sheet for OpenRouter & co?

Especially in this age where if you put AI in your company's mission statement you're drowned in money.

Let's hold off on calling something "cheap" until the external financing money runs out and the actual numbers are revealed AND audited.

> yes it's cheap.

When running toy models that do basically 0 of what regular people expect from state of the art LLMs, sure.

Running Apache is cheap. Running Google search isn't. They both serve web pages.

OpenRouter is the router. We don't care about their financials, the point is that you can buy inference via them for $x/token on a variety of models for a variety of providers. Those are businesses not propped up by SV VC dreams, just hosting plus compute and their costs.

Running a local LLM isn't a mainstream normal thing to do, sure but saying it's "basically 0 of what regular people expect from state of the art LLMs" is lazily dismissing evidence because it contradicts your beliefs. It does work, and it works at a level somewhere above "basically zero" for nerds who are willing and able to set it up for themselves today. The comparison isn't Apache, it's ElasticSearch. It's not Apache cheap and simple, but it's also not Google Spanner expensive.

> Those are businesses not propped up by SV VC dreams, just hosting plus compute and their costs.

Again, you have no way of knowing this. During a bubble a myriad small companies nobody hears about get funded for millions and billions.

Also, plenty of startups max out the founder credit cards and then they go bankrupt.

Let's revisit this discussion and see if 10% of all the companies in OpenRouter are around in 2030.

> is lazily dismissing evidence because it contradicts your beliefs

No, it contradicts my experience. The models you can run on 128GB of VRAM/unified RAM (which is realistically the maximum a regular person can buy) are basically bad compared to Anthropic/OpenAI.

More than that hardware prices spike like crazy (and even if consumers could afford them, the hardware itself is basically a huge DYI project).

Let alone the fact that regular users need to run other things on their system so can't dedicate absolutely everything to the LLM.

Again, the financials of these businesses are at best unproven and at worst critically unsound.

They're not that bad, and the point is you can run interference on 128 GB VRAM, so we have a floor on the cost of inference.

I'm sure the industry will consolidate by 2030, as industry always does.

Except for movie pass startups are marginally profitable on whatever they sell.
Inference is cheap as long as you have 1000 data centers full of GPUs. The only crux is you have neither.
But you have to keep training