Every metric becomes a target (Goodhart's law). Also, the plural of anecdote is not data.
The subjective anecdotes from HN users matter because they are not data and are much harder to game. Not impossible to game, always be aware of users with low karma, but more difficult than gaming a benchmark.
But they’re not at all meaningful after a certain point, because they’re not even attempting to explain why it works well for them. It’s just noise like this.
If somebody gives a longer explanation of why they like or dislike it that would just be more noise it doesn't really matter that much beyond their general opinion about it
I make Claude, Codex and Gemini review each other's design plan and implementation. Each always found a lot of things the others missed...until Fable 5 came out. Whatever plan or code Fable 5 comes up with, now it's very hard for Codex and Gemini to find any serious hole in it.
It's not a brand new project, but a project I've been working on and off for the last half year. I was trying to add a major feature which required some big refactoring of the current structure of the code. With the previous models I'd expect many rounds of reviews and debates between the AI agents. But with Fable 5, there's basically no debate, Codex and Gemini basically approved immediately. :)