|
|
|
|
|
by jll29
4 days ago
|
|
I don't know who downvoted the parent or why, but it's a fair question IMHO. The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic.
A proper methodology would ask each question 20 times and calculate the mean correctness across experiments. The reason is that the temperature parameter introduces random behavior. |
|