Hacker News new | ask | show | jobs
by cma 19 days ago
The scaling with reasoning models is more and more with things like verifiable rewards (coding and math), in line with bitter lesson and also Sutton invented lots of modern RL.