Hacker News new | ask | show | jobs
by nlarew 33 days ago
The frontier labs are not "fine-tuning", they're doing massive scale RL post-training