Y
Hacker News
new
|
ask
|
show
|
jobs
by
anon373839
16 days ago
I had roughly the same performance on M3 Max 64: ~30 tokens/sec. Which is not terrible, but with the latest Lightning MTP optimization I am getting ~100 tokens/sec with Qwen 3.6 35B-A3B.
1 comments
ch_sm
16 days ago
WHAT. that‘s amazing. thanks so much for sharing, i‘ve been looking for ways to speed a3b up for days. It‘s 11pm here but i‘ll try this right now
link
anon373839
16 days ago
I’m using the new oQe quants in case that matters.
link
ch_sm
16 days ago
i see! would you mind sharing your model string?
link