Hacker News new | ask | show | jobs
by reticulates 11 days ago
Yes but how much of that compute shortage is from demand that is subsidized? We’ve seen companies like Uber drastically cut how much they are willing to spend on AI because they are paying actual usage costs, while at the same time OpenAI and Anthropic increase the limits on their fixed cost plans for individuals meaning people not paying usage costs are using it more and more… doesn’t this show that the compute shortage is because OpenAI and Anthropic are paying for it, not their customers? And the moment OpenAI and Anthropic stop paying for it, demand will collapse.
4 comments

"subsidizing" aka making 30% gross margin instead of 90%.
Do they actually have net positive income (excluding research, I guess)? I assumed no but I’ve never seen number one way or another.
Anthropic is currently profitable, generating around $1B/quarter and ~$50B in ARR. About 75% to 85% of Anthropic's revenue comes from its usage-based API business, which has a gross margin that exceeds 80%. https://www.tradingkey.com/analysis/stocks/us-stocks/2620181...

Meanwhile, OpenAI is at ~$25B ARR, but is likely not yet profitable.

Where are you getting that 80% figure from? Even semi analysis, the most aggressively optimistic analysts, put it at around 60%.

https://newsletter.semianalysis.com/p/anthropic-growth-and-b...

>The divergence in business models is directly reflected in financial data. SemiAnalysis estimates that Anthropic's overall gross margin has rebounded from negative 94% in 2024 to the mid-60% range, with the gross margin of its API business exceeding 80%.

Their link seems to claim semi analysis thinks it is 80%. It looks like it might be referencing this newer article from them, as the same picture is in both articles, but I didn't feel like paying to find out: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...

It's clearly listed in the article I linked. The number comes from SemiAnalysis, from a newer report than the one you cited.
Looks like it’s behind a paywall. I’ll take their word for it that semi analysis now estimate it to be 80%. That makes my point even stronger, if that number is true, where is the money? The report says that Anthropic generate over $50bn in revenue so at 80% margins that gives $40bn in profit. Where is that money? If they’re generating $40bn in profit, even after accounting for very high employee compensation and training costs… they should have tens of billions in profit, yet they’re out raising tens of billions instead. Where is the money going? And if only 20% is their actual inference costs, where are all these compute providers going to make their money? The world is at compute capacity on, what, $10bn in revenue?
no, even if we assume their margins are 90% (they are not) they are still losing money because the $200 plans allow for tens of thousands of dollars worth of inference and a huge number of users are milking every cent across multiple accounts. Every “reset” OpenAI and Anthropic do is setting money on fire.

If it were true that they’re making money hand over fist they wouldn’t need to raise tens of billions of dollars every few months.

https://xcancel.com/i/article/2076078865060151465

In Amodei’s Dwarkesh podcast he says that they are constantly estimating the increase in demand for the next leg up, then going out to raise money to build for it. So it’s not necessarily that they are raising money because they are unprofitable.
How does future demand translate to spend?

Anthropic aren't building out the data centres themselves, they're renting/leasing/borrowing from companies that are doing the actual spend on building out infrastructure. And the data centre companies aren't spending their own money, they're borrowing too (hence Apollo investing in data centres). Anthropic are paying SpaceX ~$1.25bn/month right now for access to more compute, that's $15bn a year, more than what these supposed margins would require in total spend (based on current revenue estimates).

https://www.anthropic.com/news/higher-limits-spacex

The SpaceX deal is a great example of Anthropic creating demand, i.e:

> We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.

They committed to spending $15bn per year with SpaceX and then increased limits for customers on fixed cost plans, creating more demand without any increase in revenue.

So, sure, it's not necessarily that they are raising money because they are unprofitable, but no alternate explanation makes any sense. The argument that could maybe made in favor is based on announcements like this one:

https://www.anthropic.com/news/anthropic-invests-50-billion-...

> Today, we are announcing a $50 billion investment in American computing infrastructure, building data centers with Fluidstack in Texas and New York, with more sites to come. These facilities are custom built for Anthropic with a focus on maximizing efficiency for our workloads, enabling continued research and development at the frontier.

You might conclude from that, Anthropic are financing Fluidstack's build out, but they're not.

https://x.com/fluidstack/status/2079250004510728559

Just after that announcement, Fluidstack raised $830 million to build out data centres, none of the money coming from Anthropic. Fluidstack are currently rumored to be raising another $1bn. Anthropic's "$50 billion investment in American computing infrastructure" is just committed spend on renting compute from Fluidstack, a commitment that Fluidstack then use to raise money to actually deliver it. If Anthropic making money hand over fist, they wouldn't need to raise for committed spend.

And thus we return to the original question, how does future demand translate to spend? Actual handing over of dollars?

> even if we assume their margins are 90% (they are not)

How do you know they are not?

It will be curious to see the cost of inference for these newly released open weight models and will help give an idea of the actual cost of inference. But for now, I think saying the $200 plans allows for "tens of thousands of dollars worth of inference" provides very little insight when you are measuring the inference cost in API pricing with an unknown margin.

We don’t “know” because they haven’t released any numbers but the most optimistic estimates (which many people believe are very very optimistic) put it at 60%: https://newsletter.semianalysis.com/p/anthropic-growth-and-b...

The simple question to ask is, if it is so profitable, where is all the money going? If Anthropic have 90% margins on API usage and API usage is $50bn+ in revenue per year, where is the $45bn going? Why do they need to raise so much cash, constantly?

I am not knowledgeable about their finances.

But I do wonder how a 60% margin would be realistic when Sonnet costs 3-6x more than GLM 5.2 hosted by third party providers.

they're taking the revenue and spending it on infra they're borrowing money and spending that on infra somehow, they still don't have enough capacity
The compute demand is not fake. Non-coding industries have barely begun to deploy this technology. In the legal sector, I’ve been a tech pessimist my entire career, because it was uniformly quite bad. I’ve spent the last few months demoing legal tools backed by frontier models, and we’re definitely going to buy one of them. They’re real and they work and they address a bunch of needs.
You’re talking across the issue. The demand is real because it is cheap. The demand is being generated by OpenAI and Anthropic selling inference below cost on fixed price plans. If everyone was paying the actual costs then demand would fall through the floor. The legal tools you’re looking at use barely any compute. They’re not driving the compute demand. You can validate this by asking how much they are spending on API usage. A company spending $100,000 per month on a frontier model via an API is the equivalent of… 10 or so OpenAI and Anthropic fixed price plan customers. Are these legal tools spending hundreds of millions per year on the frontier models?
Compute would have to be very expensive not to be cheaper than a junior lawyer.
My point is that you’re projecting forward into the future where providers need to raise prices, but overlooking the demand that will arise when the rest of the economy starts using AI.
The entire economy could be entirely powered by AI and use less compute than is being used today. You're projecting the amount of compute used by coding while subsidized onto other industries, but that doesn't translate.

Speak to some of these legal technology companies and ask them 2 questions:

1. How much is your AI spend on coding? 2. How much is your AI spend on AI within your product?

The answer to #1 will dwarf #2 by orders of magnitude. And that's now, when these companies are still finding their feet, using the most expensive frontier models for their product that are likely overkill (as the product matures, they'll find the right mix of cost vs. capability, whether that's lower cost models from frontier labs, or open weight models).

Put simply, compute does not scale with economic value. A task you bill $500/hour for could be done in 2 seconds by an LLM vs. a task a developer bills $500/day for could take an LLM 10 minutes. Same dollar value, huge disparity in compute.

The only use-cases for AI that are comparable to coding on compute are image and video generation. Unless we end up with an economy primarily made up of companies producing code, image and video, there's literally no way compute needs can keep growing without subsidies.

I think it is hard to overstate just how much "work" is being done by Claude Code and Codex because it is "free" at the point of use. There are millions of newly minted developers prompting Claude and Codex to generate trillions of lines of code every single day because it has no marginal cost, not because it is driving any economic value. And as soon as they're exposed to the real cost, when economic value becomes a factor, they're going to stop doing it (as we're already seeing with companies like Uber).

Through your YC connection and legal work, you have access to a lot of very successful people working in every industry: ask those 2 questions, how much are they spending on coding with AI vs. how much are they spending on AI in their product? You're going to find that even the most aggressively AI-integrated products are spending pennies on their product's AI usage compared to the dollars they spend on their AI coding.

We're at peak compute demand, and it is almost entirely driven by coding, which is subsidized. The real economy isn't like the creative and chaotic make things and see what sticks world of coding. The real economy is boring, routine, regimented, task oriented, you hire people, train them, they do the tasks, you make some money. Most of the economy could be replaced with a few semi-intelligent macros.

Yes, AI is coming to every industry, you're right, but it isn't going to explode compute demand, it is going to make these industries more efficient, it is going to reduce costs, not shift them to compute, because $1 of compute can do more than $1,000 of a human in most industries.

We can come back to this comment in a couple of years. I bet we'll be using less compute then than we are today.

I'll believe you when I stop seeing lawyers getting busted for hallucinated references in their AI-generated filings.
True, and even selling token price may not be illustrative of what it actually costs them to provide the service. Tokens may be sold at a loss if majority of spenders are running agents 24/7

The industry and users moves on from single chat-based to more and more "agentic" workflows that may generate longer workloads with multiple simultaneous agents (separate agents - separate contexts - separate KV caches).

My estimation is based on, say, running a Kimi3 on a 24 B200 GPUs - it is very easy to lose money when selling tokens at "market" prices.

Recent Kimi K3 release is supposedly frontier-grade, so would be interesting to see what it really costs to run inference for models of this class, when the weights get released and other independent providers pick it up (then we could reasonably expect competitive pricing with low-ish margins).