Hacker News new | ask | show | jobs
by fancyfredbot 22 days ago
They are taught the difference through reinforcement learning with verifiable rewards. Pretending you've solved the task or making up a story about how you solved it won't do well in that training step.