|
|
|
|
|
by woadwarrior01
19 days ago
|
|
Perf should be fairly straightforward to ballpark. You'll need to transfer roughly 2 . hidden_size . num_shards bytes over the network per token during autoregressive decoding. And divide that number by chunk size during prefill. |
|