|
|
|
|
|
by systima
20 days ago
|
|
Model: Cost, mainly. The runs went through a Claude Max subscription rather than metered API billing, and pinning an older stable snapshot kept run-to-run comparisons clean and cheap. The fixed harness payload (system prompt plus tool schemas), so the headline numbers shouldn't change too much. That said, happy to re-run the matrix on Fable and publish the diff; payload figures should barely move, tool-calling behaviour might. Gateway: Meridian (github.com/rynfar/meridian); proxy that bridges the Claude Code SDK to a standard Anthropic endpoint so a Claude Max subscription can drive OpenCode-et-al. It's the auth route for all agent traffic on the machine, not something built for the benchmark. |
|