Y
Hacker News
new
|
ask
|
show
|
jobs
by
_davide_
21 days ago
I'm writing my own inference engine for Strix Halo and the same model. I already have 30%+ performance plus a more graceful decay over long contexts; that said, their point stands: memory bandwidth is what you really want.