Hacker News new | ask | show | jobs
by NitpickLawyer 565 days ago
> So once they go down a path they can’t properly backtrack.

That's what the specific training in o1 / r1 / qwq are addressing. The model outputs things like "i need to ... > thought 1 > ... > wait that's wrong > i need to go back > thought 2 > ... etc