|
|
|
|
|
by verdverm
18 days ago
|
|
That's what happens when you quant too hard. I'm working on quant strats and evals for the same underlying qwen 27b models. When I saw 27b on a phone, I thought not fitting, big phone, or aggressive quant. NVFP4 still takes 27G before KV cache. |
|