I think there's two halves to the conversation: which models have more weights and which models are better than the other ones listed. I think this was about the latter part. There are plenty of smaller models these days which knock the socks off older models 10x the size.
We are talking about the top tier of open weights models - GLM-5.2, DeepSeek V4, Qwen3, Kimi K2. The ones ranking on leaderboards and giving frontier labs a run for their money. Gemma may have its uses but it is not in that conversation.
I find it easy enough to interpret as talking to "a better non-Chinese model since Llama 3", even if it doesn't compete with the current large Chinese models mentioned before that. Again, they aren't saying you're wrong about large models - they are adding that Llama 3 was separately surpassed by smaller non-Chinese models during the wait.
British-American with much of the research happening in London. I don't know if it's known what team specifically worked on which of the Gemmas, I'd suspect it's a healthy mix of multiple satellites across the globe, but the contribution by the UK Deepmind parts likely was significant, certainly not so small that it'd justify downing a regular question.