|
|
|
|
|
by HarHarVeryFunny
1 day ago
|
|
It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though. |
|
Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.
No part of this pipeline is fixed in stone.