|
|
|
|
|
by petesergeant
25 days ago
|
|
My dissertation used a Likert scale with ChatGPT-3 era LLMs, and it was both internally consistent over 5-6 runs on a given statement, and consistent with human raters. I don't think you can bat it away as simply "doesn't work" |
|