|
|
|
|
|
by gpugreg
1 day ago
|
|
Notably, MXFP4 was introduced at the (much less costly) supervised fine-tuning stage after pretraining, so the number of B200/B300 GPUs could be relatively small in comparison to the number of H200 GPUs used during pretraining (or maybe not, who knows). (Kimi K3 tech report section 4.1.1 https://arxiv.org/pdf/2607.24653) |
|