Y
Hacker News
new
|
ask
|
show
|
jobs
by
andai
14 days ago
Results seem mostly noise to me. One eval per model, in a large problem space (i.e. a problem which requires many attempts to solve well).
1 comments
couAUIA
14 days ago
Yes I agree, but I actually did a lot more runs, with different prompts, different times ect... And each time /goal had a small or insignificant impact
link