|
|
|
|
|
by lhl
27 days ago
|
|
You got me curious, so I made a little harness comparison to my model test suite: Model Adapter Suite Score Passed Tasks
--------------------------------- ------------ -------------- ------ ------ -----
local/ornith-1.0-35b little_coder aider_polyglot 36.0% 81/225 225
local/ornith-1.0-35b pi_devstack aider_polyglot 39.6% 89/225 225
local/ornith-1.0-35b pi_vanilla aider_polyglot 32.0% 72/225 225
Little Code does a little better than raw Pi, although maybe not better than my personal Pi setup: https://github.com/lhl/devstack |
|