Related, I was given access to mimo-v2.5-ultraspeed, which is amazing. This is now my expectation for speed, it’s fast enough for me to stay mentally engaged rather than getting distracted waiting for the agent to churn.
The -spark variant of GPT was a ton of fun indeed, such a shame it's so dumb though so really hard to rely on. If you were to compare the quality of mimo-v2.5-ultraspeed with anything from OpenAI/Anthropic, where would it be placed ~more or less in your view?
Is it the same quality as base mimo v2.5, or different? I've been enjoying regular mimo v2.5 quite a bit via opencode, if ultraspeed provides the same quality, that's crazy.
I'm not a regular v2.5 user, so I don't really know. But given the TileRT team write-up says the entire network gets quantized to FP8 (and experts get quantized to FP4)[0], I'm assuming there's at least a modest drop in quality.