|
|
|
|
|
by realusername
41 days ago
|
|
There's also a lot of benchmark trickery going on, it's becoming harder to see how the latest models really improved. The top models also seem to have inconsistent performance depending on the time of day and how far we are from the next release. |
|
Even with minor automation I feel like I can watch OpenAI and Anthropic engineers fiddling in real-time. Tuesdays behaviour changes by Thursday, 10AMs production isn’t possible at 11:30AM. Nutty.