The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
If you run a model on slower hardware are you getting more experience? Surely its a factor of model output reviewed and not human time.