|
|
|
|
|
by m00dy
1 day ago
|
|
I would want to see three things before drawing strong conclusions: End-to-end tokens/sec and cost on realistic coding agent trajectories, including tool outputs and retries, not isolated decode benchmarks. Cache hit rates and prefill cost for branching, multi-turn sessions. Router-load distributions after post-training, where expert collapse or specialization problems often show up. |
|