Hacker News new | ask | show | jobs
by singingtoday 20 days ago
GPT was the worst on the Rubik's cube
4 comments

GPT-5.5 isn't really a fair comparison to Fable or Opus. GPT-5.5-Pro would be a better comparison (and yes, I know how much more expensive it is).

I really wish they'd thrown in something like GLM-5.2 into the comparison.

Grok did not render anything, they had to prompt it again.
They gave grok 2nd try on that one. Now one shotting a webpage is a dumb metric but if that is what you are testing it did the worse.
GPT doesn't render but does actually work, Grok does not fully work Two scrambles in a row and the thing is broken.