|
|
|
|
|
by henryrobbins00
6 days ago
|
|
Some additional details about these results that didn't make it into the main post: - The OpenATP "standard provers" were used; see docs [5] for model / harness configuration details - Time and cost are function of effort level, which may lead to unfair comparison across provers - FATE-X excludes task 10 since claude and grok hit session limits - FATE-X excludes leanstral and aristotle due to temporary endpoint failures - Deepseek's FATE-X accuracy is corrected from 2 to 3 due to verifier bug (now fixed) - 2 FATE-X deepseek misses are sorry-free, but rely on native_decide - Claude's FATE-X miss is due to a failed delegation to a background subagent - All costs come from underlying CLI, except codex which uses pricing table |
|