|
|
|
|
|
by ul5255
18 days ago
|
|
I’m staring at this comment for a while now: With 3ms latency combined per token, wouldn’t that mean (1 / latency) = 333 token/s for the theoretical upper bound? I’m not trying to nitpick, just curious if I misunderstand something. |
|
33 tps max token generation speed would be for 10ms of network latency in the above example.