Y
Hacker News
new
|
ask
|
show
|
jobs
by
singingtoday
20 days ago
GPT was the worst on the Rubik's cube
4 comments
trollbridge
20 days ago
GPT-5.5 isn't really a fair comparison to Fable or Opus. GPT-5.5-Pro would be a better comparison (and yes, I know how much more expensive it is).
I really wish they'd thrown in something like GLM-5.2 into the comparison.
link
GaggiX
20 days ago
Grok did not render anything, they had to prompt it again.
link
cuvinny
20 days ago
They gave grok 2nd try on that one. Now one shotting a webpage is a dumb metric but if that is what you are testing it did the worse.
link
sceptic123
20 days ago
GPT doesn't render but does actually work, Grok does not fully work Two scrambles in a row and the thing is broken.
link
I really wish they'd thrown in something like GLM-5.2 into the comparison.