devtool
SWE-bench Verified no longer measures frontier coding capabilities
Abandoned benchmark tool no longer suitable for evaluating cutting-edge coding capabilities
3.1Overall
Utility3
Onboarding5
Craft5
Niche fit2
Longevity1
Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.
Good for
Only for academic research on historical model performance
Not for
Any scenario requiring evaluation of current AI coding capabilities
Project description
openai.com
Alternatives
HumanEvalMBPPCodeContests