Hacker News new | ask | show | jobs
by anon373839 29 days ago
Can you mention what inference stack you're using? I've tried MTP several times with that model and it always seems to significantly cut my token generation speed from ~60 tokens/sec to ~40 (M3 Max).
1 comments

(see above reply to myself) I misattributed the gain from dflash - it was dflash-mlx library + mlx model, not dflash itself giving me the speedup.