|
|
|
|
|
by not_a_bot_4sho
13 days ago
|
|
If you're not doing *at least* say 100 iterations (thousands are preferred!!), you do not have enough data to draw any stable conclusions. Interestingly enough, using an LLM-as-judge is a great way to approach things like this at scale but you do need to invest in some Cohen's Kappa or Fleiss' Kappa understanding which means putting a human in the driver seat to evaluate the effectiveness of your non-human judge. Absent of that, it's just another case of human-centipede but with LLMs. |
|
What does "better" even mean there?