Hacker News new | ask | show | jobs
by QuercusMax 26 days ago
As long as LLMs don't understand the different between regurgitating facts and making up stories, they're going to necessarily be limited.
1 comments

They are taught the difference through reinforcement learning with verifiable rewards. Pretending you've solved the task or making up a story about how you solved it won't do well in that training step.