Y
Hacker News
new
|
ask
|
show
|
jobs
by
maxignol
29 days ago
Did not seem to find how much tokens per second he achieved with this setup ?
1 comments
aetherspawn
29 days ago
80 tok/s which is kind of a lot for GLM. My experience running 80 tok/s on other LLM is that it ~seems faster than cloud inference, but that obviously depends what you use, in my case ChatGPT.
link