| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by m_ke 117 days ago
	this is mostly because RLVR is driving all of the recent gains, and you can continue improving the model by running it longer (+ adding new tasks / verifiers) so we'll keep seeing more frequent flag planting checkpoint releases to not allow anyone to be able to claim SOTA for too long