Hacker News new | ask | show | jobs
by toomuchtodo 5 days ago
My primary role is cybersecurity in a regulated entity in a regulated industry, I am highly confident it is straightforward to do so based on work accomplished in only a couple of weeks. Stand up a router, stand up a Kubernetes cluster if you don't have one, stand up the necessary VMs and compute for serving inference. Two pizza team, in my experience.

Customers can switch (although we can argue the speed and pain of doing so), and the speed at which they do will be a function of cost efficiency and demonstrable value (imho). A recent example of this is Broadcom and VMware [1], for example. When motivated, it can be done. If there is no objective, measured value being delivered, the spend will be cut. If the value delivered is measured, it will be enabled at a lower cost through cost optimization measures (ie self hosting) [2].

This is all to say: there is no moat, the revenue of inference providers is volatile and not assured in any measure. Caveat emptor.

[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

[2] Microsoft considers replacing ChatGPT and Claude with Kimi K3 to save $600M - https://news.ycombinator.com/item?id=49022984 - July 2026

4 comments

There's a difference between "It's straightforward to do" and "I can convince a customer to sign a contract allowing us to do it."

If the second one were as easy as the first, I wouldn't have to be online at 9:00 to deploy stuff to prod tonight; the team in India would handle it. But customers write into the contracts that only US-based employees interact with prod systems. No amount of cajoling will get them to change their minds; they have data sovereignty, international telecommunications treaties, and HIPAA compliance to worry about. So I'll be pressing buttons tonight.

Could you swap out Anthropic or OpenAI or Google or whoever's models for Kimi? Yes. They're like other software these days, they're modular. What isn't modular is regulatory and geopolitical concern.

Meanwhile, in real companies, you have to wait 2 months or more to access an API endpoint in preprod.

To setup a cross business kubernetes cluster will take 2 years with unknown results.

On Cloud, in Switzerland, you need to call Microsoft when you need new resources, so much for agility and minute infrastructure provisioning, and I heard the same for AWS.

> Meanwhile, in real companies, you have to wait 2 months or more to access an API endpoint in preprod.

> To setup a cross business kubernetes cluster will take 2 years with unknown results.

Do you seriously believe those times will not go down 95% if the CEO pushes for it to get done yesterday because it will save the company millions in expenses?

You get it. Given sufficient incentives, processes and systems become potentially more malleable, and hard requirements can become optional. Speed is a function of appetite, will, and resources.
> Stand up a router.

it has to be some amazing router and while the models are open-weights, the knowhow to run them efficiently surely is not?

Good luck explaining to an exec that the locally hosted Chinese model definitely doesn't have a backdoor or hidden trained-in intentions.

Meanwhile the cost/benefit analysis doesn't move much even if you are paying 2x for tokens, and you don't need anything on prem.