Hacker News new | ask | show | jobs
by LaurensBER 18 days ago
It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet.

It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is.

Surely, from a purely economic perspective it would be wiser to keep this proprietary and benefit from the increased API traffic?

6 comments

What Chinese firms are doing makes perfect sense from the commercial perspective actually because they understand how a classic commoditization spiral works. The reality is that models themselves are general commodities and there's just not enough difference between them. A company can get ahead of others by a few months, but then the rest quickly close the gap. It's a really low margin business because there's no way to differentiate yourself.

Chinese companies know that there's no profit in general purpose models in the long run, and they're treating models as shared infrastructure akin to Linux. They're amortizing the cost of research by keeping models open, and rapidly closing the gap and driving prices towards the marginal cost of inference. The money is going to be in customization niches. Companies will charge to tune models for specific use cases and charge support for that. There's also going to be money at the bottom for hardware vendors making chips and memory. But the middle tier of generic LLMs is seeing involution where there's relentless competition driving profits towards the bottom.

The Chinese labs incentives is to run inference for the world, because inference can run on the homegrown Chinese chips (giving them a guaranteed market for their hardware) and they have cheap plentiful power.

The US frontier labs have an incentive to do deals with large firms to act like a contract research organization, taking royalties on creations/discoveries. Alex Karp called this out in his rant ("Why charge for tokens, take a %") and he's basically right about this.

Yeah China has a huge (and growing) advantage in power generation. And they have been looking to break into high-end chip manufacturing. The latter has high cost of entry and needs large scale to become viable. AI inference on own hardware would allow them to bootstrap the chip demand. Both for memory and accelerators.
I expect that US companies will ultimately angle for long term government contracts similarly to companies like Raytheon. There's basically unlimited money available, and they don't have to worry about competition from China here.
Their game? Sell me tokens instead of me buying them from an American lab for a higher price.

Publishing open weights gives me more confidence in the model, and ironically makes me less anxious about making sure I can replace the cloud usage with a local alternative. Whereas I’m very nervous right now with relying on 5.6-Sol - what if they triple the price, nerf it, etc.?

> Publishing open weights gives me more confidence in the model

Why? It's not like you can audit weights like you can with code.

> what if they triple the price, nerf it, etc.?

What if an open weights infra provider does that? What's the difference?

Because I can run Qwen 3.6 or DeepSeek V4 until the end of time if I want to? The model is on HuggingFace; anyone can download it. I have Qwen and Gemma on my laptop right now if I want to use them, even if I decide to go be a hermit who never interacts with the outside world again.
I will concede a use case for hermits.
Or travel. Even in the developed world you can be without internet or slow internet. I have Gemma 4 E4B on my phone that can process audio, image, and text if I have need to.
Ok but we're talking about the concern that a provider can 3x or nerf a model. I'm pointing out that open weights providers can do that too, and your only recourse is to run it yourself at way more than 3x the price.
You might not be able to audit the weights, but there are people with the skill set to do it.

If providers decide to jack the price, open weights lets you find a new provider without losing your fine tunes and having to re-do workflows, etc like you would if you switched off a frontier lab model.

you keep the model. it's never deprecated. with closed ai, you are forced into a new more expensive model every few months. if an open model infra provider does that, you simply switch to another one. it's not in their interest to do that.
That's not really true though, providers are deprecating models and I have at least 10 emails to prove it.
Nothing stops you from downloading the model and hosting it on a cloud virtual machine
Common sense does, but other than that I suppose you're right.
A provider deprecating a model doesn't mean the .gguf file disappears from my computer.
hell yeha bro, I still rock qwen 0.1
When the weights are publically available (and open to use), then there's a free market for hosting that model at least. That's not independence for me, but it's lack of vendor lock-in.

For some of the open models, there's a list of 20-30 providers of the same model on openrouter for example, as an example of the supply.

> the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations

At Xiaomi, MiMo is now led by Luo Fuli. She is a former Alibaba & DeepSeek employee: https://newsen.pku.edu.cn/news_events/news/people/15385.html (https://archive.vn/I8Pmu) / https://e.vnexpress.net/news/tech/personalities/who-is-luo-f... (https://archive.vn/sb3B6)

Don't know if it is due to Luo, but it is striking how similar performance & pricing of the models, DeepSeek v4 Pro & MiMo v2.5 Pro, is.

Christ, she had such a nicer vibe than Altman or Amodel.
The bet could be that they’ll ultimately be able to sell hardware capable enough of running local models comfortably.
If China is able to undercut Nvidia on high performance local AI hardware, they will pop the AI bubble in a matter of days.

They wouldn't even need to make something equivalent to the latest hardware. A Chinese RTX 3090 equivalent would be enough.

Standard commoditize your complement.
its a governement mandate that states that AI research must be open source ,that's one benefit of communism
There is no such mandate. ByteDance keeps their models closed. So does iFlyTek. Qwen Max is closed as well.
But they are not groundbreaking. They are simply copies of other similar Chinese models

The mandate is worded differently from what I said

What is the literal wording of this "mandate" then?