Hacker News new | ask | show | jobs
by apical_dendrite 26 days ago
The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
1 comments

However, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."
I didn't get why they mentioned that one specifically. Is there any particular relationship between Blueprint-bench and Vendor-bench?
Both benchmarks are made by the same people.