Hacker News new | ask | show | jobs
by noosphr 4 days ago
There is more than one workload in AI. Inference for llms is memory constrained on even a single card. Training for llms is memory constrained on the level of racks. In both cases you hardly ever see more than 40% of advertised flops used.
1 comments

So the question is: why are they designing AI cards so unbalanced?