|
|
|
|
|
by hmate9
11 days ago
|
|
This model is not for builders and engineers. DeepSWE score of 49% is behind gpt 5.4 and muse spark. It's clearly intended to be an efficient model for google gemini usage. What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models. |
|
Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.
I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.
https://alignment.openai.com/measuring-reward-seeking/