|
|
|
|
|
by noosphr
4 days ago
|
|
There is more than one workload in AI. Inference for llms is memory constrained on even a single card. Training for llms is memory constrained on the level of racks. In both cases you hardly ever see more than 40% of advertised flops used. |
|