Hacker News new | ask | show | jobs
by m00dy 2 days ago
There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.
3 comments

They cant push it too low. The license agreement it is released under wont allow it.

> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

Yeah, this license is a lot different to Kimi K2.7 Code or Kimi 2.6, previously it stated:

"Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2.7 Code" on the user interface of such product or service."

Although as you mentioned, now there is a revenue cap before custom agreements must take place. Currently there is 7 providers on OpenRouter for Kimi K3 and they have all the exact same price unfortunately.

I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.

I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)

I would be surprised, considering that OpenRouter requires disclosing the quantization and shows automatic benchmarks to compare between providers for the same model.
The latter is a joke
I saw that it runs GPQA Diamond and TAU-Bench Airline and shows the results over a 32 day rolling average.

Other than that they track Tool call error rate and Structured output error rate.

I only discovered this today, and it seems like a good idea. What are the problems in practice?

As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.