Y
Hacker News
new
|
ask
|
show
|
jobs
by
fancyfredbot
22 days ago
They are taught the difference through reinforcement learning with verifiable rewards. Pretending you've solved the task or making up a story about how you solved it won't do well in that training step.