← Tool Radar
dataset

APIEval-20

API testing benchmark dataset with limited coverage and uncertain community adoption

4.8Overall
Utility5
Onboarding6
Craft5
Niche fit4
Longevity4

Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.

Good for

Research teams needing quick evaluation of AI agents' API calling capabilities

Not for

Enterprises requiring production-grade API test coverage

Project description

An open benchmark for AI agents that test APIs Discussion | Link

Alternatives

Postman CollectionsSwagger/OpenAPI test suites
Visit siteEvaluated 2026-05-09

Similar tools