Hacker News new | ask | show | jobs
by sgk284 20 days ago
If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/

I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

4 comments

I want this with smaller models as well like Gemma 4 or Qwen 3.6
Awesome - will work on getting those in.
It would be nice to see how "long" -- e.g. how many self-turns/tool calls/etc the prompt takes to resolve as well.

I know that models like Gemma4-e4b will take longer self-turns but IBM's Granite models will take shorter self-turns in exchange for more tool calls.

Having 'local runable' to compare would be awesome. For example I have a 48G MacBook.
It'd be interesting to add Nemotron, which is quite popular on Spark alongside Qwen 3.6.
Really nice site! From your experience, what’s your go-to model for nice storefronts?
This is impressive and, I think, complements well whatever benchmark is the hottest right now.
Impressive, specially the amount of models used for the comparison.