|
|
|
|
|
by upbeat_general
1 day ago
|
|
This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute. Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count. No part of this pipeline is fixed in stone. |
|
I think the word "distillation" needs to be used a bit more selectively here. If their pre-training run was complete before Fable was released that implies that ZERO Fable data went into the base model. Perhaps the timeline allows for a few weeks at best of incremental post-training on some limited amount of Fable data, but calling this "distillation" seems a bit dramatic especially given the redacted outputs that would have been available. A more factual speculation would just be that they may have had time to post-train using a limited amount of Fable output in some fashion (LLM as judge? SFT? Who knows ...).