... and it was not SOTA at the time of release. Gemini 3.1 Pro was previewed on 19 February 2026. It just barely beat GPT-5.3 and was roughly equal with Sonnet 4.6. Better according to some, worse according to others. And that lasted about 10 days (until GPT-5.4 came out which was also not a huge jump). And these are large averages. On coding or terminal Gemini 3.1 Pro was not close to Opus 4.6. Also it only matches the GLM 5 open model, more or less.
Last time Google had a "everybody agrees" SOTA model was Gemini 3 in November 2025, it beat GPT-5.1, it really was better and held it for a month.
Currently Google's best model doesn't match open models, in any category (performance, price or speed), in fact the last 4 open weight champions all beat Google's best model.
> Last time Google had a "everybody agrees" SOTA model was Gemini 3 in November 2025
I have no opinion on whether this is true, but "It's been 6 months since Google had the best model in the world" seems a rather weak criticism!
It's anyways been clear for a long time that people are finding value at all sorts of different model sizes and price points, and that Pareto frontier and cost-to-complete task are more important than who benchamaxxed who.
If a model as strong as Gemini 3.6 Flash(!) had been released a year ago, then everyone would be falling over themselves calling it AGI - it is extremely capable, and free usage in the chat app is essentially unlimited.
Really? Weak criticism given that in that time Google was beaten by OPEN models? China is no longer behind Google in AI, Google is now catching up to China and open models. Every linux shop can give their customers better AI performance than Google can, at a cheaper cost than Google's flagship model (either $100k up front and marginal cost only, or about half the cost of Gemini 3.6 Flash).
At least OpenAI and Anthropic can still tell their customers their best model is better than some Finnish student can setup in the customers' basement, even if the difference is pretty small now. But they are better. Google is not.
Surely that's a major change in the situation. "America" is losing the advantage (although even the Chinese models are losing, see next paragraph) but Google is a big step behind 6 other players and 3 open models, so let's just say, Google is effectively behind everyone else that matters. Worse: one of the models that beats Google's best model was trained on Chinese ASICs (that aren't even 2nm. And very few I might add, couple thousand. Google's next model ... has a LOT to prove)
Even Chinese models are losing their advantage. By which I mean that if you check how far Qwen 3.6 27B is behind trillion-parameter models, it's less than a year, down from easily 3 years. If that continues to go down, even Chinese trillion parameter models will become hard to justify (and Qwen 3.8 is being cooked up as we speak, I mean we don't even know if a new 27B is in the works or not, but everyone's very excited)
Has Google even being trying to keep up with frontier sized models, and for that matter why should they?! Even if they want to, what's the rush? Someone from Google just tweeted today that they just recently STARTED their "Gemini 4" (whatever that may be) pre-training run... In the meantime they just today announced financial results with cloud revenue up 82% YOY ... they seem to be doing just fine.
Let's put it this way. The Q&A on the earnings call yesterday, or about 50% of it went like this:
Bank 1> What is Google going to do about not having a SOTA model?
Sundar> Everyone uses Gemini 3.6 Flash anyway. We do that internally as well.
Bank 2> What is Google going to do about not having a SOTA model?
Sundar> Flash is great. We've started our most ambitious training run with Gemini 4.
Bank 3> What is Google going to do about not having a SOTA model?
Seriously, like 4 banks in a row, same principle. I'll look up the exact wording in the transcript, but for now it's not available and I'm not listening to the 20 minutes of that repeat again.
As for your question: Google is training a new major version of Gemini, so yes, Google is trying to get a SOTA model. They've just not succeeded for a while.
OK, so Wall St apparently want Google to get into the (as yet unprofitable) frontier LLM competition. But why? To help any of Google's existing businesses? (doesn't make sense - the smaller much more efficient Gemini 3.6 Flash is much better suited to that).
Now, Google DeepMind is still very much pursuing AGI, and do see an LLM as being a component of that, so that's at least one area where it might make sense for Google to invest in a frontier LLM, but otherwise ???
Why aren't Wall St clamoring for Google to get back into robotics to compete with Tesla, or maybe to get into the rocket business to compete in the "data centers in space" business, or maybe get into the dry cake mix business to compete with Sara Lee ? ...
I don't agree with you on what the term SOTA means. Gemini Pro 3.1 was the only frontier model that had native video input at launch. It was certainly SOTA at some tasks.
That's my point. It wasn't a clear win. It did some things no other model did. And it sucked at coding and terminal (as in nowhere near SOTA). It wasn't a clear win.
Last time Google had a "everybody agrees" SOTA model was Gemini 3 in November 2025, it beat GPT-5.1, it really was better and held it for a month.
Currently Google's best model doesn't match open models, in any category (performance, price or speed), in fact the last 4 open weight champions all beat Google's best model.
Here's a plot of the evolution: https://www.reddit.com/r/LocalLLaMA/comments/1v20g29/kimik3_...