Hacker News new | ask | show | jobs
by Zaheer 443 days ago
Impressive how well Grok performs in these tests. Grok feels 'underrated' in terms of how much other models (gemini, llama, etc) are in the news.
2 comments

you can't download grok's weights to run locally
how is that relevant here?
it helps explain why theres' less people talking about them than gemini or llama?

less people using them.

You can't download Gemini's weights either, so it's not relevant as a comparison against Gemini.

I think the actually-relevant issue here is that until last month there wasn't API access for Grok 3, so no one could test or benchmark it, and you couldn't integrate it into tools that you might want to use it with. They only allowed Grok 2 in their API, and Grok 2 was a pretty bad model.

lol sorry mixed them up w gemma3 which feels like the open lesser cousin to gemini 2.5/2.0 models
I can guarantee you none of my friends (not in tech) use “downloading weights” as an input to select an LLM application.
isn't chatgpt the most used or most popular model?
Yes OpenAI has a first-mover advantage and Claude seems to be close as a second player with their closed models too, open weights is not a requirement for success but in an already crowded market (grok's prospect) their preposition isn't competing neither with top tier closed models nor the maybe lesser-capable but more-available battle-tested freely available to run locally open ones
It's not.

Also, only one out of the ten models benchmarked have open weights, so I'm not sure what GP is arguing for.

> in terms of how much other models (gemini, llama, etc) are in the news.

not talking about TFA or benchmarks but the news coverage/user sentiment ...

I am amazed Gemini did as well as it appears.

Gemini frequently avoids discussing health problems, which likely hurt its scores. My guess is any censorship was considered a fail.