|
|
|
|
|
by sparsesignal
22 days ago
|
|
Nice paper. They predict the outcome of edits up to ~25 steps before the agent makes them. Decodable doesn't mean causal, as the authors note, but a cheap probe that flags doomed trajectories early could save a lot of wasted agent compute. We clearly still have a lot to learn about what these models represent internally. |
|