|
|
|
|
|
by crs_gentleman
29 days ago
|
|
We tried this, and it works :) https://arxiv.org/abs/2508.06111 You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned. Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve. We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised. |
|