Hacker News new | ask | show | jobs
by wolttam 20 days ago
It's the native DeepSeek V4 Flash quant which is released in MXFP4. You have enough RAM left over with 2 sparks for ~2M tokens of working KV cache.