|
|
|
|
|
by jasonjmcghee
27 days ago
|
|
I usually get mlx versions of models and didn't realize I was using non-mlx - misattributing the performance boost to speculative decoding. With dflash-mlx library, with dflash disabled 3.6-35b-a3b mlx model I'm getting ~60 t/s on low power and >100 t/s on high power. Compared with previous message being official qwen huggingface release (non-mlx) using lm studio. |
|