Hacker News new | ask | show | jobs
by 1337h4xx 34 days ago
The model's reward function is its own reward. No need to reinvent the wheel.