|
|
|
|
|
by can1357
5 days ago
|
|
I'd love nothing more than measure 3), but quality is really hard to assess IMO as even with the frontier models I find myself disagreeing on what is good and maintainable code; so that disqualifies LLM-as-judge, which even at it's best is somewhat RNG and costly. + I don't see this replacing anyones workflow to make it clear. This is mostly for fully autonomous agents that are given a task & are asked to execute. For instance, subagents are a perfect example where you can get the aforementioned savings for free, even if you don't run a "software factory" per-se, as you generally do not interact with them. |
|