| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by zozbot234 35 days ago
	True, ARC is mostly an artificial "human-like AGI" benchmark that doesn't really reflect any plausible workload. Very different from things like Humanity's Last Exam that reflect real-world knowledge and are now getting closer and closer to saturation even with open models.