Hacker News new | ask | show | jobs
by siquick 22 days ago
Likert scale just doesn't work in LLM evals. My idea of 3/5 is different from your 3 and definitely different from an non-deterministic system's 3.
2 comments

That's what the rubric is for; it reduces the problem to NLP, which is the LLM's forte. The more objective you can make your rubric, the better.
My dissertation used a Likert scale with ChatGPT-3 era LLMs, and it was both internally consistent over 5-6 runs on a given statement, and consistent with human raters. I don't think you can bat it away as simply "doesn't work"