Some groups are baking models into silicone, Deepmind has an example, it gets 18,000 tokens/sec on Llama 3.1, not sure about parameter size
While some other groups are baking silicone into models :)
but yes misplaced e