|
|
|
|
|
by b-zee
39 days ago
|
|
Mentioned directly under the table: > Note GPT 5.5 Pro is at the top of the leaderboard only because it blew through $100 budget after only completing four cases, so 2/4 is 50%. And, a couple of other results, both Qwen models, are skewed upward in the detect % ranking because of failure to complete all cases. |
|