Hacker News new | ask | show | jobs
by jasonjmcghee 27 days ago
I usually get mlx versions of models and didn't realize I was using non-mlx - misattributing the performance boost to speculative decoding.

With dflash-mlx library, with dflash disabled 3.6-35b-a3b mlx model I'm getting ~60 t/s on low power and >100 t/s on high power.

Compared with previous message being official qwen huggingface release (non-mlx) using lm studio.