Hacker News new | ask | show | jobs
by arendtio 22 days ago
I also like to take a look at https://cursor.com/cursorbench

While in the past months Composer 2.5 was a lot better than I had expected a year ago, and the GPT 5.6 family does a good job in terms of cost for performance, I wonder why nobody is talking about Grok 4.5 high? Those numbers look very convincing to me.

1 comments

> * Grok 4.5 has an advantage on CursorBench: an earlier snapshot of the Cursor codebase was unintentionally included in training. The exact score impact is unclear. That data has been removed for future models. For a rundown of third-party benchmark scores, see the Grok 4.5 launch blog.

I don't know about those numbers, even assuming this was by mistake :)