Hacker News new | ask | show | jobs
by maxgashkov 6 days ago
I don't really get why you need handoff if your score is accurate. If it is, and it is low for a given response, just let the harness re-run the prompt with a different seed until the score is high enough. If this approach doesn't work, your score is most likely garbage.
1 comments

Smaller models in production can produce quite interesting outputs.