| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by NitpickLawyer 166 days ago
	swe-rebench is a pretty good indicator. They take "new" tasks every month and test the models on those. For the open models it's a good indicator of task performance since the tasks are collected after the models are released. A bit tricky on evaluating API based models, but it's the best concept yet.