| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by scrollop 225 days ago
	Alright so we have more benchmarks including hallucinations and flash doesn't do well with that, though generally it beats gemini 3 pro and GPT 5.1 thinking and gpt 5.2 thinking xhigh (but then, sonnet, grok, opus, gemini and 5.1 beat 5.2 xhigh) - everything. Crazy. https://artificialanalysis.ai/evaluations/omniscience

1 comments

On your Omniscience-Index vs. Cost graph, I think your Gemini 3 pro & flash models might be swapped.