dataset
APIEval-20
API testing benchmark dataset with limited coverage and uncertain community adoption
4.8Overall
Utility5
Onboarding6
Craft5
Niche fit4
Longevity4
Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.
Good for
Research teams needing quick evaluation of AI agents' API calling capabilities
Not for
Enterprises requiring production-grade API test coverage
Project description
An open benchmark for AI agents that test APIs Discussion | Link
Alternatives
Postman CollectionsSwagger/OpenAPI test suites