| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by NitpickLawyer 565 days ago
	> So once they go down a path they can’t properly backtrack. That's what the specific training in o1 / r1 / qwq are addressing. The model outputs things like "i need to ... > thought 1 > ... > wait that's wrong > i need to go back > thought 2 > ... etc