|
|
|
|
|
by saberience
19 days ago
|
|
"On Agents’ Last Exam (opens in a new window), an evaluation of long-running professional workflows across 55 fields, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost. " Some pretty big claims and results! Excited to see how it feels during usage. I use Fable and 5.5 extensively and I still find both have a place in my toolkit, i.e. Fable IS good but it isn't perfect, and it's still better to play them off against each other. I have Fable and 5.5 write plans and have them adversarially review each other's plans. Having this amount of competition in the coding model space is good for all of us. |
|