A cheaper model can look fine per request, but one weak answer may create another call or a human review step. That seems like the hardest cost to capture.
True, but human review isn't visible in the logs, hmm, however I suppose retry patterns can, but frugon doesn't currently track this, so we can't quantify this yet.
Frugon analyzes your existing logs offline on cost and quality tier. The "looks fine per request" is exactly the reason why --judge exists.
Earlier, cyanydeez inspired the consideration of a metric "effective cost per judged success", which could also answer this point.
Try it on your logs, and tell me if anything falls through the cracks.
Frugon analyzes your existing logs offline on cost and quality tier. The "looks fine per request" is exactly the reason why --judge exists.
Earlier, cyanydeez inspired the consideration of a metric "effective cost per judged success", which could also answer this point.
Try it on your logs, and tell me if anything falls through the cracks.