Hacker News new | ask | show | jobs
by verdverm 21 days ago
What quant size are you using? Just picked up my second spark and a cable and wanted to try this one out for fun. I generally run a bunch of models for different purposes (easy / bulk tasks) and use APIs for harder tasks.
1 comments

It's the native DeepSeek V4 Flash quant which is released in MXFP4. You have enough RAM left over with 2 sparks for ~2M tokens of working KV cache.