|
|
|
|
|
by xienze
16 hours ago
|
|
> There are huge token efficiency/bloat differences between agents while working on the same tasks, using the same model, in the same environment Be careful here. Remember these are non-deterministic models at the end of the day, and even with everything being "the same" you can have two runs where the same model, same harness, same tools can arrive at the same conclusion through a wildly different sequence of events. |
|
I will add more tasks (esp longer ones) and think more about grading, the current tasks were easy to grade because the desired outcomes are well specced but I will also look into more open ended tasks and how to grade those
thank you!