Y
Hacker News
new
|
ask
|
show
|
jobs
by
verdverm
21 days ago
What quant size are you using? Just picked up my second spark and a cable and wanted to try this one out for fun. I generally run a bunch of models for different purposes (easy / bulk tasks) and use APIs for harder tasks.
1 comments
wolttam
20 days ago
It's the native DeepSeek V4 Flash quant which is released in MXFP4. You have enough RAM left over with 2 sparks for ~2M tokens of working KV cache.
link