Hacker News new | ask | show | jobs
by hmate9 11 days ago
This model is not for builders and engineers. DeepSWE score of 49% is behind gpt 5.4 and muse spark. It's clearly intended to be an efficient model for google gemini usage.

What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models.

3 comments

Presumably 3.6 Flash is primarily meant to serve their own needs for the Gemini chat app, voice app (which Sergey Brin says he uses a lot in the car on the way to work), and for their search "AI Assistant".

Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.

I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.

https://alignment.openai.com/measuring-reward-seeking/

Who tested it on DeepSWE?

Edit: Oh it's in the other link

https://blog.google/innovation-and-ai/models-and-research/ge...

Everyone wants to announce as late as possible (i.e., last) to chart the highest. Google is in a position financially to take a hit for these last few months.