Hacker News new | ask | show | jobs
by anon373839 16 days ago
I had roughly the same performance on M3 Max 64: ~30 tokens/sec. Which is not terrible, but with the latest Lightning MTP optimization I am getting ~100 tokens/sec with Qwen 3.6 35B-A3B.
1 comments

WHAT. that‘s amazing. thanks so much for sharing, i‘ve been looking for ways to speed a3b up for days. It‘s 11pm here but i‘ll try this right now
I’m using the new oQe quants in case that matters.
i see! would you mind sharing your model string?