|
|
|
|
|
by fweimer
28 days ago
|
|
I think the decode phase of inference typically uses local compute resources poorly due to the very small batch size. If you can run many inference tasks in parallel, this will make local inference more competitive to centralized inference, not less. |
|