I want a super fast LLM that is Opus 4.6+, like, in ability.
For inference workloads, it makes a lot more sense to optimize for prefill/ttft before maxing out memory bandwidth.