|
|
|
|
|
by toomuchtodo
5 days ago
|
|
My primary role is cybersecurity in a regulated entity in a regulated industry, I am highly confident it is straightforward to do so based on work accomplished in only a couple of weeks. Stand up a router, stand up a Kubernetes cluster if you don't have one, stand up the necessary VMs and compute for serving inference. Two pizza team, in my experience. Customers can switch (although we can argue the speed and pain of doing so), and the speed at which they do will be a function of cost efficiency and demonstrable value (imho). A recent example of this is Broadcom and VMware [1], for example. When motivated, it can be done. If there is no objective, measured value being delivered, the spend will be cut. If the value delivered is measured, it will be enabled at a lower cost through cost optimization measures (ie self hosting) [2]. This is all to say: there is no moat, the revenue of inference providers is volatile and not assured in any measure. Caveat emptor. [1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... [2] Microsoft considers replacing ChatGPT and Claude with Kimi K3 to save $600M - https://news.ycombinator.com/item?id=49022984 - July 2026 |
|
If the second one were as easy as the first, I wouldn't have to be online at 9:00 to deploy stuff to prod tonight; the team in India would handle it. But customers write into the contracts that only US-based employees interact with prod systems. No amount of cajoling will get them to change their minds; they have data sovereignty, international telecommunications treaties, and HIPAA compliance to worry about. So I'll be pressing buttons tonight.
Could you swap out Anthropic or OpenAI or Google or whoever's models for Kimi? Yes. They're like other software these days, they're modular. What isn't modular is regulatory and geopolitical concern.