|
|
|
|
|
by atlex2
1 day ago
|
|
What was the architecture of your router? If it was based on GRPO/RL, it would be interesting to hear why your router performance capped. I think the truth is that it's not an efficient cost cutting method. Your router has to be at least as 'smart' as all the but the smartest of your models (models do poorly when asked 'is this a task you're well suited to'), and that means you're caching multiple prompt histories including kv-filling/prefix caching on your expensive router model. Most of the time, not super great for savings. |
|