Hacker News new | ask | show | jobs
by zozbot234 17 days ago
It's a very nice summary of KV cache storage requirements for the leading open weights models, but ultimately it still shows the DeepSeek V4 series winning by a huge margin on that front. And DeepSeek actually does quite well on needle-in-haystack context recall benchmark so it seems like the outstanding context compression it has isn't even impacting smarts all that much. Keep in mind that KV cache memory requirements usually has a direct impact on how much you can parallelize inference on both large-scale platforms and lower-end edge compute, so this is a very important metric.