Y
Hacker News
new
|
ask
|
show
|
jobs
by
mi_lk
26 days ago
Cursor: Find me another benchmark where Composer 2.5 is a top 10 frontier coding model
1 comments
leerob
26 days ago
(I work at Cursor) We score well on Terminal-Bench and SWE-bench Multilingual. DeepSWE, not so great yet, as it's more for very long-horizon tasks. We're planning to include more public benchmarks in our next model release.
link