Hacker News new | ask | show | jobs
by larrysalibra 10 days ago
Ben's article "distills" down to 2 reasons that US frontier labs shouldn't be "afraid":

1. US frontier lab unit economics are better 2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.

For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.

For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.

And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.

And every improvement in model capability makes it increasingly easier to make your own tools.

Wrote more on this in a blog post that has an earlier HN discussion: https://news.ycombinator.com/item?id=48982061

Direct link: https://larrysalibra.com/ben-thompson-is-wrong-us-frontier-l...

6 comments

He’s glossing over the reason they are not: 90% profit margin of Nvidia. Power is only a small part, single digit, it will eventually matter but does not really today.

What is the cost of AI? The single largest ingredient is Nvidia profit margin.

Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.

Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.

Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)

> Power is only a small part, single digit, it will eventually matter but does not really today.

Sorta yes, sorta no.

A single 5090 consumes 450W - at Californian energy prices of $0.38 per kWh that's $0.17 per hour. And the card itself costs $4100 on amazon. So after 2.75 years running at full power 24/7 you'll have spent more on electricity than on the card. I would have thought most data centres being built today would have a design life longer than 3 years.

Of course you can throttle the cards to ~300W without losing too much performance. But also you need more than a single 24GB card to run most modern LLMs.

DC has wider margins than the 5090. by someone elses rough numbers (https://www.spheron.network/blog/gb300-nvl72-vs-gb200-nvl72-...) for GB300 rack:

> 132kW

> ~$3.7-4M

So about 300x the 5090's power but 1000x the price. Roughly 9 years for electricity to exceed price at $0.38 and datacenters will show up in areas with cheaper power than CA.

Anyone seriously building out AI infrastructure I presume is paying nowhere near $.38/kWh which is extortionate. Utility scale solar is closer to $.02-.03/kWh, then maybe around ~$.10/kWh for natural gas peaker plants.
Utility scale power price varies widely by location and exact time of day.

LLM's aren't very latency sensitive and can therefore move to wherever power is cheapest.

Right now that's places next to aluminium smelters (which also like very cheap electricity 90+% of the time).

GPU design life is commonly cited as about 3 years. The main driver is not physical degradation of the cards but the expectation they'll become obsolete with newer cards.
Also add electricity 50% on top for cooling the DC.
Cost of electricity isn’t a long term advantage in my opinion. Private companies will figure it out.

What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.

I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.

Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.

I'm a model nomad, using whatever solved my last problem the best and where it makes the most sense to start my next work in.

However with the latest models Fable, Kimi K3, 5.6, it's getting to a point where I sometimes forget what model I am on without noticing a difference. And once I realize it because something may not be exactly like I expected it I won't switch for that work either because I don't want to invalidate the cache.

For the next work I will do there is maybe a 50/50 chance to remember to switch the model before I start.

That's not what I would call stickiness towards a certain provider.

You didn't say what kind of problems you solve with AI. It matters a lot if you are doing HTML versus C++, for example.
In no order of importance:

  - Refactoring a 13 year old in-house vacation rental booking system ( python/turbogears )
  - Backend development for our VR fitness game ( flask/python )
  - Unity development on our VR fitness game ( C#/Unity )
  - VR game development experiments ( Godot/GDScript )
  - Standalone SLAM localization service ( C++ )
  - Audio analysis ( python/pytorch )
  - Virtual display with Viture display glasses ( C )
  - Reverse engineering a library I am using for another project ( ghidra -> C - no MCP yet, that's something I am looking forward to )
  - Public facing website rebuilding for the booking system above ( PHP/JS )
  - Generative 3D environments for our VR fitness game ( python )
  - Wireless camera/IMU based tracker for the SLAM system ( C )
Once I've dug in with a specific model into a problem I tend to stick to that because I have a feeling what it will do and how well it works, but when I start a new thing I usually use whatever the model was last set to.
Wow, now we're talking :)
Seems like most popular harnesses, including codex and Claude code, support Agent Skills (an open spec for skill formatting/ organization): https://agentskills.io/clients

Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)

you can simply switch to z.ai/GLM-5.2 inside Claude Code by settings env variables in .claude/settings.json
> Private companies will figure it out.

Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy

You really have to look at energy/capita and how much energy is embedded in exports. The gross numbers are misleading. The US wasn't building new electrical generation capacity because it didn't need it and there was no market for it (caveats apply, but in a broad sense this is the major reason). Now that the market exists the question is how much can the US actually bring online and how rapidly, which is a real challenge after decades of degrowth politics used to justify slash and burn consumption of the industrial base.

AI is really all about electricity. AI could be completely fake and yield zero value whatsoever and the US would do exactly what it is doing now because the AI bubble is what creates the market for building new electrical generation capacity, which is needed for re-industrialization. Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.

> The US wasn't building new electrical generation capacity because it didn't need it

You need to understand that the US added 40.3 GW of energy capacity in 2023. 70 to 80 represents a dramatic growth (nearly doubling) in new capacity.

Also China is just getting started. 8 out of the 12 nuclear power stations opening worldwide in 2027 are Chinese https://world-nuclear.org/information-library/current-and-fu...

If AI is "all about electricity" as you say then the US was barely ever in the running in the first place. There's no possible way the US could compete any time in the next 2 decades (putting aside a dramatic shakeup like war)

> Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.

Lol people make up the funniest theories when a political idea they don't like is gaining popularity. At least you didn't blame Russian bots

Oh, the UK and Europe castrated themselves and want others to follow the example. They still don't have made the connection between "I give up ability" and i get attacked by a proxxy opponent by those i gave ability up too.

They do not want to life in the world that is and thats going to be, but in the past and the world they green ideology promised. Reality denial be a addictive poison.

That last para is a novel idea. I wonder if there's any evidence to support/undermine it though?
The concepts of industrial reserve capacity and using dual-use consumer goods to subsidize military production capacity are well known and widely practiced historically in the US. China adopted this strategy from the US, and the US conveniently forgot about it for a few decades in order to justify selling off the industrial base to China, but at least based on public documents like the published U.S. National Security Strategy, I would assess with high probability that this is explicitly recognized and being followed now.

AI is not fake and it does work, but what I am saying is that from a pure systemic analysis perspective, you can do the numbers, and even if AI was complete fugazi, the benefits you get from the electrical generation capacity, and the ability to fund it through private markets, which bypasses Congress, and locks in commercial contracts (often with foreign governments) which will be almost impossible politically to reverse, would still make it optimal from a strategic perspective. That is my calculation, and to the extent that it is correct, I would assume that the US Military's strategic planning apparatus would arrive at the same conclusion.

AI compute has some unique characteristics that make it especially useful for grid management. Moving consumer compute to the cloud means that the electrical use of that compute can be centrally managed. In an emergency, you can cut electrical use for consumer AI by 50% or more, because chips run more efficiently at lower power, and you can shift workloads onto quantized models, reduce resolution for video output, etc, to reduce compute, which leads to minor service degradation but not interruption. AI datacenters are also adding massive amounts of battery storage capacity, which is an additional grid buffer. For every GW in capacity added by hyperscalers that is creating a dispatchable reserve capacity of 50% under completely normal circumstances (hyperscalers do this internally to optimize their own costs) and then that number goes up depending on the scale and duration of the emergency.

They are not stopping either.
I'd really love to see the evidence on this!
> 1. US frontier lab unit economics are better

That's not generally true, since there is generally still much reliance on NVIDIA. The true low cost providers are Google with their TPU and vertically optimized stack, and Amazon with Trainium. However, Google does not have their own frontier model, and Anthropic (who are partially served by Amazon) are also paying a premium for extra NVIDIA-based capacity from SpaceX, maybe soon from Meta too.

I don't know how the economics of domestic Chinese Huawei-based clouds (no NVIDIA) compares to the west, but since serving cost is mostly hardware depreciation and to a lesser extent electricity, they are not necessarily at a disadvantage (Ascend 950 costs roughly 50% of an NVIDIA H100), and more to the point it is irrelevant when considering US commercial use that is more likely to be using Chinese open weights models from US providers served on NVIDIA based hardware.

I think the real significance of Chinese frontier models being open weight is that it takes development cost amortization out of the US-based serving cost, while the US AI labs can't afford to do this. The US labs therefore need to reduce development spending to remain price competitive. The Chinese companies are of course still making money from the Chinese market, whether by selling API access or by other business models such as Ziphu making 75% of it's total revenue by selling services to Chinese customers who are running their models on-prem due to the Chinese apparently being very concerned about data privacy.

More basically, production cost matters only if inference is priced at commodity prices. That's not what VC's signed up for, which is rent-seeking.
In a corporate setting yes Opencode all the way. However in a non corporate setting I am getting $3000 of api usage a month for $100 at Anthropic and only use open code for the smallest cheapest tasks
I think calling opencode the trend is naive. This not what is being run on company time.
OpenCode is absolutely used on company time.
Shhhhhh...! OpenCode only runs on authorized machines by responsible employees following company policy ;)