|
|
|
|
|
by NitpickLawyer
565 days ago
|
|
> So once they go down a path they can’t properly backtrack. That's what the specific training in o1 / r1 / qwq are addressing. The model outputs things like "i need to ... > thought 1 > ... > wait that's wrong > i need to go back > thought 2 > ... etc |
|