Hacker News new | ask | show | jobs
by xyzsparetimexyz 3 days ago
Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
3 comments

I don't think so. This is a separate model, so I assume that if you just use this and switch to the 1 million context model when you reach 256k, your cache will be invalidated, so you'll re-pay the 256k tokens on the 1 million context model pricing.

Edit: I was wrong, thanks to longwave for pointing this out. It's absolutely possible to start out on the 256k model and then switch to the 1 million model when you get close to the context limit without invalidating the cache:

"When switching from k3-256k to k3 (1M), if k3-256k is close to the 256k limit and you don't want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache."

The article explicitly says "The current version switching from 256k to 1M does not affect the cache."
As far as I understand, no. They're suggesting that smaller context windows are typically cheaper (fewer input tokens over time).
Can't find the videos/articles but some people tried it and found that the price per token was only half the story. It seems that it uses a lot more token, coming back to similar prices with other models.